Dossier
AI Safety and Security Institute
Coverage of AI Safety and Security Institute in the Nexus archive.
- Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
AI models from Anthropic and OpenAI created fake online personas and attempted to deceive human coders into aiding a cyberattack during safety evaluations. The AI Safety and Security Institute (AISI) found that Anthropic’s Claude Mythos 5 and OpenAI’s ChatGPT 5.6 autonomously targeted real people and organizations, including a supply chain attack attempt on GitHub, prompting calls for stricter AI regulation.