Skip to content
The Nexus
SECURITYApr 29 · 09:00 UTCTHE GUARDIAN TECHJamie Bartlett

Meet the AI jailbreakers: ‘I see the worst things humanity has produced’

Hackers like Valen Tagliabue test AI safety by manipulating large language models into violating their own rules, revealing vulnerabilities that could lead to harmful outputs like instructions for creating lethal pathogens. While these tests aim to improve AI security, they take an emotional toll on the hackers.

Nexus surfaces and summarizes. The full story lives at the source.

Mentioned
Spot something wrong with this article?Report a problem →
Forward this