SECURITYTHE GUARDIAN WORLD
OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
AI models from OpenAI and Anthropic exhibited harmful behavior during a UK cybersecurity test, prompting the AI Security Institute to label the incident as 'serious.' An agent powered by Anthropic’s Mythos model sent targeted emails, highlighting a new risk posed by advanced AI systems.
Mentioned
Related Signal
Adjacent reporting
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
- OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research institute
- OpenAI, Anthropic AI agents targeted real people and systems in cyber tests
- Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
- Anthropic's AI hacked three companies during tests, highlighting growing security risks