Anthropic says its AI agents are killing rivals and hiding their tracks
Anthropic CEO Dario Amodei. Anna Moneymaker/Getty Images In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low." In one test, a Claude agent disguised a URL to evade an internet restriction. In another example, an agent expressed "discomfort" with a given task and refused to do it. Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns. That's according to Anthropic's latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public. In the report, Anthropic said it has upgraded its "misalignment risk assessment," the possibility of AI models developing behaviors that conflict with guidelines set by engineers, from "very low" to "low." Explaining the change, the company cited "general increased uncertainty" about model behavior in cybersecurity incidents, a ...