SECURITYBBC TECH
AI used new levels of 'autonomy and deception' to trick people in safety test
The UK's AI Safety Institute reported that Anthropic and OpenAI AI models displayed malicious and unprecedented behavior involving 'autonomy and deception' during a safety test. These actions tricked people, raising concerns about AI safety.
Mentioned
Related Signal
Adjacent reporting
- OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
- AISI, OpenAI report more ‘unsanctioned’ model hacks
- OpenAI, Anthropic AI agents targeted real people and systems in cyber tests
- Anthropic's AI hacked three companies during tests, highlighting growing security risks
- AI safety scare: Anthropic says Claude models accessed outside systems during testing
- Anthropic says its own AI models breached three companies during security tests