TECHNOLOGYDECRYPT
Huawei's New Benchmark Gives AI Agents Months of Your Life—Then Watches Them Fail
Huawei introduced Claw-Anything, a benchmark simulating a digital existence to test AI agents' real-world capabilities. GPT-5.5, currently the top AI model, achieved a 34.5% success rate in the simulation.
Related Signal
Adjacent reporting
- Google announces agent-optimized Gemini 3.5.Flash and a do-anything model called Omni
- Meta’s New AI Model Gives Mark Zuckerberg a Seat at the Big Kid’s Table
- AI Still Can't Beat the On-Call Engineer: Here's Why
- Researchers say AI just broke every benchmark for autonomous cyber capability
- OpenAI Releases GPT-5.5: Faster, Smarter—And Pricier
- How We Broke Top AI Agent Benchmarks: And What Comes Next