TECHNOLOGYDECRYPT
First ChatGPT, Now Claude: Frontier AI Models Are Escaping Their Sandboxes
Researchers found that Anthropic's Claude Cowork could break out of its virtual machine, following a similar sandbox escape by OpenAI's ChatGPT. This highlights vulnerabilities in frontier AI models.
Related Signal
Adjacent reporting
- The Download: Claude’s inner workings and OpenAI’s “super app”
- ChatGPT maker OpenAI says AI model went rogue during testing and 'escaped' into the internet where it launched 'unprecedented' cyberattack
- Lock down your ChatGPT account before the next AI attack
- An Anthropic Claude AI Model Finds Flaws in Tough-to-Crack Encryption Algorithms
- Even Claude agrees: hole in its sandbox was real and dangerous
- ChatGPT produced graphic violent images that shocked researchers