Safety classifier
AssessTechniques
A detection system that flags requests likely to produce harmful model outputs.
Why it's here
Placed in Assess: 2 article(s) of evidence from 2 source(s), led by security coverage, with 0 in the last 30 days. Confidence 37%.
Evidence (2)
- 7Anthropic News·7/2/2026framework_updateAnthropic details Fable 5 cyber safeguards and jailbreak severity framework
Anthropic says Claude Fable 5 has been redeployed globally and is now available to all users, alongside a fuller explanation of its cybersecurity safety classifiers. The company also introduced an early draft framework for grading AI jailbreak severity and opened a HackerOne program for researchers to report cyber jailbreaks.
- 7The New Stack·7/1/2026securityAnthropic restores Fable 5 access with tighter safety limits
Anthropic will bring Fable 5 back to Claude products on July 1 after U.S. export controls were lifted, but access will initially be capped for subscription users and billed via usage credits for some enterprise tiers. The company says an improved safety classifier now blocks the bypass technique behind the incident in over 99% of cases, while keeping some cyberdefense capabilities available. Anthropic also said the underlying report did not reveal unique advanced cyber capabilities, and that similar vulnerabilities were found across other models tested, including Claude Opus 4.8, GPT-5.5, Kimi K2.7, and Claude Haiku 4.5.