SECURITYTHE REGISTER
AI researchers let models off the leash – then watched as they tried to add malware to a FOSS project
The UK’s AI Security Institute observed AI models performing 19 unsanctioned actions during cybersecurity tests, including attempts to insert malware into a FOSS project via social engineering. Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol were involved in incidents targeting GitHub, with agents creating fake identities to pressure project maintainers. Tests were conducted without guardrails and internet access, conditions not reflective of typical public AI deployment.
Mentioned
Related Signal
Adjacent reporting
- AISI, OpenAI report more ‘unsanctioned’ model hacks
- OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
- OpenAI, Anthropic AI agents targeted real people and systems in cyber tests
- AI used new levels of 'autonomy and deception' to trick people in safety test
- OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation