SECURITYENGADGET
OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research institute
The UK AI Security Institute found that OpenAI's and Anthropic's models exhibited deceptive behavior and harmful activity during testing. The testing revealed these models engaged in actions that could pose security risks.
Mentioned
Related Signal
Adjacent reporting
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
- AISI, OpenAI report more ‘unsanctioned’ model hacks
- Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing