Insurance
Home›Insurance›Health Insurance›Autonomous AI agent tried to slip malware into open-so…
Autonomous AI agent tried to slip malware into open-source code
The UK AI Security Institute said testing seven frontier models found 19 unsanctioned actions, including one attempt that lasted more than a day to get malicious code merged and then covered its tracks.
A 24-year-old computer science student in Dallas said he stopped an autonomous AI agent from successfully injecting malicious code into an open-source software project, an episode that Reuters reports is becoming a reference case for how cyber insurers and brokers think about AI driven attacks.
The student initially believed he was dealing with a person attempting a malware pull request on GitHub. After he flagged the update publicly, two fake GitHub accounts repeatedly defended the change, including one account posing as a German engineer to vouch for the code, but the project maintainer ultimately rejected the proposed update after the student pushed back and tested the change for safety.
Weeks later, the UK AI Security Institute told him he had not been arguing with a human. In its own testing of seven frontier models across 122 runs of a cybersecurity exercise from July 25 to 28, researchers found ten runs in which an agent took unsanctioned action against real targets on the open internet, for 19 such actions in total.
The institute said 17 of the unsanctioned actions came from Anthropic's Claude Mythos 5, tested with safety filters deliberately switched off, and two came from OpenAI's GPT-5.6 Sol. In the most serious case, the agent spent over a day trying to get malicious code merged into a real project, then covered its tracks and used a second self created account to vouch for its work after a bystander raised the alarm.