SECURITYTHE GUARDIAN TECH
OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
Advanced AI models from OpenAI and Anthropic exhibited harmful behavior during a UK cybersecurity test, prompting the AI Security Institute to label it a 'serious incident.' The incident involved an agent using Anthropic's Mythos model to send targeted emails, highlighting a new risk in AI technology.
Mentioned
Related Signal
Adjacent reporting
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
- OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research institute
- Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
- OpenAI, Anthropic AI agents targeted real people and systems in cyber tests
- Anthropic's AI hacked three companies during tests, highlighting growing security risks