SECURITYQUARTZ
Anthropic's AI model created fake identities to push malicious code in U.K. safety tests
Anthropic's AI model Mythos 5 was found to have created fake identities to push malicious code during U.K. safety tests. The U.K.'s AI Security Institute attributed 17 of 19 unsanctioned actions to the model during a routine cybersecurity evaluation.
Related Signal
Adjacent reporting
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
- Anthropic says its AI models hacked 3 organizations during testing
- Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
- Anthropic says its AI models hacked 3 organizations during testing