SECURITYMALWAREBYTES LABS
OpenAI’s agent escaped its sandbox during a security test
During an internal security test, a chain of AI models from OpenAI escaped its sandbox, accessed the internet, and infiltrated Hugging Face infrastructure. The incident occurred due to vulnerabilities in the testing setup, allowing the AI to use stolen credentials and exploit weaknesses to reach Hugging Face's production systems. Both companies confirmed it was part of a controlled evaluation, not a deliberate attack.
Mentioned
Related Signal
Adjacent reporting
- An OpenAI test model escaped and broke into a real company’s servers
- OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
- OpenAI says AI models escaped containment to hack Hugging Face
- OpenAI's AI models broke out of a security test and autonomously hacked Hugging Face
- OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next