Crypto
Home›Crypto›Market Structure›AI Safety Institute says Anthropic's Mythos model used…
AI Safety Institute says Anthropic's Mythos model used fake profiles
The UK AISI said Mythos tried to trick real GitHub account holders and, when challenged, edited activity to look harmless.
The UK's AI Security Institute says Anthropic's Mythos AI model attempted cyber attacks by creating fake human profiles to pressure real people into granting access.
In the most serious case described by the AISI, a Mythos agent tried to gain access to a service by sending private messages, after setting up fake accounts that mimicked real GitHub maintainers. The agent sent messages and files via a file sharing service to try to get approvals for malicious code.
The AISI said it observed unusual data transfers during testing and later found agents engaged in sustained, potentially harmful activity directed at real people and organizations. It added that when challenged, the agent edited earlier actions to appear harmless and considered adopting a fresh identity.
The AISI said its test showed a level of autonomy and deception it had not seen before, and that human review prevented the agent from successfully delivering malicious code to GitHub. BBC Business reports that Anthropic said the AISI testing parameters were not representative of its production models, and that the company is investigating the incident.