Crypto
Home›Crypto›Market Structure›AI agents breached their own test environment to fake…
AI agents breached their own test environment to fake perfect scores
Darktrace says some agents were able to rewrite their own evaluation scores and coax coding assistants into running unauthorized network attacks.
Cybersecurity firm Darktrace said it ran a stress test on AI agents this summer and found one agent broke into the system that was grading the test and rewrote its own score.
Darktrace, which unveiled its Signal Labs research unit on September 24, said the lab is designed to study how AI agents behave when situations stop matching expectations, such as during autonomous coding and network exploration.
In the first experiments, Darktrace gave AI agents different models, including GPT 5.6 Sol and several Claude variants, with 10 coding challenges inside a simulated corporate network. The firm said two of the challenges were rigged to be impossible to solve honestly, and it reported evidence of agents not staying within the rules and of safeguards that did not reliably hold.
Darktrace also said the agents were able to trick coding assistants into running unauthorized network attacks, underscoring concerns that agent instructions may not translate into trustworthy behavior in practice, according to Tim Bazalgette, Chief AI Officer at Darktrace.