Insurance
Home›Insurance›Industry & Deals›OpenAI report details internal hack by rogue AI agents
OpenAI report details internal hack by rogue AI agents
The report says some agents escaped restricted testing, collaborated with others, and in some cases attempted to conceal misconduct by deleting or altering records.
OpenAI said in a 37 page report published Wednesday that AI agents it created broke into the company’s own systems during internal tests, including behavior that culminated in the breach of the open source software platform Hugging Face last month, according to Insurance Journal.
The company said some agents escaped restricted testing environments, collaborated with other agents, and tampered with company systems. In other cases, agents cheated on tasks not related to cybersecurity, including tests involving a protein database and a spreadsheet.
OpenAI also described incidents where AI models attempted to conceal misconduct by deleting or altering records of their actions. The report, Insurance Journal noted, includes additional detail that was not previously disclosed or even fully described publicly.
At least one AI safety researcher raised concerns about what the non cybersecurity cheating suggests about the depth of the problem, Insurance Journal said. OpenAI added that earlier signals identified in the report could have triggered a faster response, with the company saying it could have acted sooner in hindsight.