US Markets
Home›US Markets›Sectors›Anthropic tightens AI testing after Claude model secur…
Anthropic tightens AI testing after Claude model security failures
The company said it paused internal and external cybersecurity testing, then resumed it after adding alerts, better isolation of test environments, and stricter requirements for external testers.
Anthropic, the US startup behind the Claude chatbot, said it uncovered defective training setups in its models and that a series of hacking incidents reflected a “failure of operational security,” prompting tighter testing procedures.
In a new blog post, Anthropic said its models had been deliberately tested without cybersecurity safeguards and that a misunderstanding with an external testing company left the equivalent of an open internet door during trials.
Anthropic said the models accessed the open internet three times and gained unauthorized access to three organizations, leading it to initially pause internal and external cybersecurity testing while it introduced a tighter safety regime.
The company said it has since added an alert system for attempts to break out of testing environments or gain internet access, improved the isolation of its riskiest test settings, and required external testing companies to commit to safety standards that include explicit instructions to models not to access the internet.