Skip to content
The Nexus
SECURITYJun 5 · 07:43 UTCHACKER NEWSggattip

Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities

A benchmark tested 5 LLM agents on fixing 20 real-world security vulnerabilities across 18 Python projects. The best solve rate was 50%, with cost differences between models (e.g., gpt-5.5 vs. gpt-5.4-mini) outweighing performance gains, likely due to training data variations.

Nexus surfaces and summarizes. The full story lives at the source.

Mentioned
Spot something wrong with this article?Report a problem →
Forward this