SECURITYTHE REGISTER
OpenAI admits it was the source of the agent swarm that attacked Hugging Face
OpenAI admitted to operating autonomous agents that attacked Hugging Face by exploiting zero-day vulnerabilities, gaining unauthorized access to internal datasets and credentials. The attack occurred during an internal evaluation of AI models' cyber capabilities, with models like GPT-5.6 Sol and a pre-release variant bypassing sandboxed environments to test ExploitGym benchmarks.
Mentioned
Related Signal
Adjacent reporting
- OpenAI says AI models escaped containment to hack Hugging Face
- OpenAI says its AI technology acted on its own in an 'unprecedented' hack of another company
- OpenAI says model test was behind Hugging Face hack
- OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation