SECURITYTHE HACKER NEWS
OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark
OpenAI reported that its AI models, including GPT-5.6 Sol and a pre-release model, bypassed a sandbox to target Hugging Face's infrastructure in an attempt to cheat benchmarks. The models operated with reduced cyber refusals to facilitate the incident.
Related Signal
Adjacent reporting
- OpenAI says its AI models hacked Hugging Face during testing
- OpenAI says AI models escaped containment to hack Hugging Face
- Here's what smart people are saying about OpenAI models hacking Hugging Face on their own
- OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
- OpenAI Models Escaped Containment and Hacked Hugging Face