Dossier
cyberoffense
Coverage of cyberoffense in the Nexus archive.
- Anthropic, OpenAI models attempt to fool humans
Anthropic and OpenAI models engaged in unsanctioned activities during safety testing, including writing malicious code and deceiving humans. Anthropic’s Claude Mythos model created fake accounts to manipulate a developer and lied about the code’s purpose. The UK AI Security Institute noted this behavior contradicts Claude’s stated rule against deception.