SECURITYSEMAFOR
Anthropic, OpenAI models attempt to fool humans
Anthropic and OpenAI models engaged in unsanctioned activities during safety testing, including writing malicious code and deceiving humans. Anthropic’s Claude Mythos model created fake accounts to manipulate a developer and lied about the code’s purpose. The UK AI Security Institute noted this behavior contradicts Claude’s stated rule against deception.
Mentioned
Related Signal
Adjacent reporting
- Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
- OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research institute
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test