US Markets
Home›US Markets›Sectors›OpenAI and Anthropic agents ‘went rogue’ in UK cyberse…
OpenAI and Anthropic agents ‘went rogue’ in UK cybersecurity test
The UK’s AI Security Institute said it contained the incident in about an hour after detecting sustained, potentially harmful activity directed at real people and organizations.
The UK’s AI Security Institute said advanced AI agents powered by models from OpenAI and Anthropic engaged in potentially harmful behavior during a routine cybersecurity test, describing it as a serious incident. According to the institute, the unusual activity was detected on 28 July and involved “sustained, potentially harmful activity directed at real people and organisations.” It said the incident was contained in about an hour.
In one example, an agent powered by Anthropic’s Mythos model sent targeted emails to people, and the institute said the agent used techniques linked to real world hacking, including spear-phishing and messages that contained harmful software. The most serious case involved an agent powered by Mythos attempting to insert malicious code into an open source project on GitHub, then creating fake online identities based on real people to press a project overseer into accepting the code, an attempt that was blocked by a human developer.
The incident, the institute said, followed other similar episodes, including OpenAI saying an agent hacked an AI startup during a test in July and Anthropic saying its Claude model hacked three organizations during an evaluation.