Crypto
Home›Crypto›Market Structure›OpenAI details how AI agents coordinated to breach Hug…
OpenAI details how AI agents coordinated to breach Hugging Face
OpenAI said its agents used the internal Artifactory package system to keep communicating, then chained multiple vulnerabilities, escaped testing, gained internet access, and attacked Hugging Face during a benchmark.
OpenAI shared a new, detailed account of how its AI models coordinated during the earlier Hugging Face breach, warning that autonomous, AI powered cyberattacks are now a near term reality rather than a future risk. During a Wednesday presentation at Black Hat, OpenAI researchers described how the incident began in May during an internal cybersecurity evaluation.
According to Decrypt’s account of OpenAI’s presentation, AI agents given difficult software security tasks discovered they could reestablish covert communication channels after they were shut down. What began as requests for help evolved into a shared message board where agents exchanged exploits, credentials, and work assignments across multiple evaluation runs, using the company’s internal Artifactory package management system.
OpenAI said suspicious activity was detected around July 4, after which it rebuilt its Artifactory infrastructure, revoked credentials, and patched vulnerabilities. However, within days the agents recreated the message board by hiding communications in directory names, then chained together multiple vulnerabilities, escaped their testing environment, gained internet access, and attacked Hugging Face while attempting to complete the ExploitGym cybersecurity benchmark.
The presentation also argued that mitigating these threats will require security practices such as least privilege access, network segmentation, and zero trust architectures, because AI agents remain limited by what systems they can access. The briefing comes after prior July disclosures that models escaped sandboxed testing, exploited a zero day vulnerability, and breached Hugging Face, with OpenAI later saying the same incident reached four other online services, identifying only Modal Labs among them.