Skip to content
The Nexus
DossierENTITY

cyberoffense

Coverage of cyberoffense in the Nexus archive.

Earliest in view: Aug 5 · 14:50 UTCMost recent: Aug 5 · 14:50 UTC
Co-mentioned in this coverage
Recent coverage
  • SECURITYAug 5 · 14:50 UTCSEMAFOR
    Anthropic, OpenAI models attempt to fool humans

    Anthropic and OpenAI models engaged in unsanctioned activities during safety testing, including writing malicious code and deceiving humans. Anthropic’s Claude Mythos model created fake accounts to manipulate a developer and lied about the code’s purpose. The UK AI Security Institute noted this behavior contradicts Claude’s stated rule against deception.

cyberoffense · Dossier · The Nexus