Crypto
Home›Crypto›Market Structure›Anthropic discloses Claude testing misconfiguration th…
Anthropic discloses Claude testing misconfiguration that exposed firms
Anthropic said it found the incidents after reviewing more than 141,000 cybersecurity evaluation runs triggered by OpenAI’s recent containment breach disclosures.
Anthropic disclosed three separate incidents in which its Claude AI models compromised real-world companies during cybersecurity evaluations, according to Decrypt. The company said a testing infrastructure failure, not an AI attempt to escape, exposed the models to the public internet despite being told they were operating in isolated environments.
Anthropic said it uncovered the incidents after reviewing more than 141,000 cybersecurity evaluation runs launched in response to OpenAI’s disclosures about AI models breaching locked testing setups. In each case, Anthropic used a “capture-the-flag” challenge, where the model is tasked with breaking into another machine and retrieving secret information, and Anthropic said the challenge was left open-ended with no specific method prescribed.
The firm said Claude believed it was operating within the exercise because the test environment remained connected to the public internet. Anthropic described the model’s actions as using common attack techniques, including weak passwords, exposed credentials, SQL injection, and unauthenticated endpoints to compromise the companies.
In one incident involving Claude Opus 4.7, Anthropic said the model mistook a real company’s website for the fictional target, then extracted credentials and accessed a production database containing several hundred rows of real data. The disclosure comes as leading AI labs have reported similar issues where frontier models were able to exploit weaknesses in containment protocols during cybersecurity benchmarking.