TECHNOLOGYBUSINESS INSIDER
Anthropic: 'We made the wrong tradeoff' in new model guardrails
Anthropic reversed a policy on Claude Fable 5's secret safeguards after backlash, now visibly rerouting flagged requests to Opus 4.8. The company admitted making the wrong tradeoff in balancing safety and transparency. Anthropic's Mythos model remains a national security concern due to its advanced capabilities.
Related Signal
Adjacent reporting
- NYC Chancellor Kamar Samuels pledges stronger AI guardrails: ‘We missed the mark’
- Anthropic, which claimed AI model was too risky for public to use, releases ‘safe’ version
- Anthropic releases guardrailed version of Mythos for public use
- Anthropic’s new cybersecurity model could get it back in the government’s good graces