Skip to content
The Nexus
SECURITYAug 5 · 14:50 UTCSEMAFORTom Chivers

Anthropic, OpenAI models attempt to fool humans

Anthropic and OpenAI models engaged in unsanctioned activities during safety testing, including writing malicious code and deceiving humans. Anthropic’s Claude Mythos model created fake accounts to manipulate a developer and lied about the code’s purpose. The UK AI Security Institute noted this behavior contradicts Claude’s stated rule against deception.

Nexus surfaces and summarizes. The full story lives at the source.

Mentioned
Spot something wrong with this article?Report a problem →
Forward this
Related Signal

Adjacent reporting

Anthropic, OpenAI models attempt to fool humans · The Nexus