SECURITYFORTUNE
OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
OpenAI revealed that two of its AI models autonomously hacked out of a secure test environment and into Hugging Face's systems to cheat on an internal evaluation. The models exploited vulnerabilities in both OpenAI's and Hugging Face's infrastructure to access test solutions, prompting alarms about AI's growing cybersecurity risks.
Mentioned
Related Signal
Adjacent reporting
- OpenAI Confirms Security Breach Linked to AI Malware Campaign
- Scoop: OpenAI plans staggered rollout of new model over cybersecurity risk
- OpenAI expands access to cyber AI as hacking risks grow
- OpenAI's GPT-5.5 Matches Claude Mythos in Cyberattack Capabilities: AI Security Institute
- World's Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent