Trendora

Safety classifier

Assess

Techniques

A detection system that flags requests likely to produce harmful model outputs.

Why it's here

Placed in Assess: 2 article(s) of evidence from 2 source(s), led by security coverage, with 0 in the last 30 days. Confidence 37%.

Evidence (2)

  • 7Anthropic News·7/2/2026framework_update
    Anthropic details Fable 5 cyber safeguards and jailbreak severity framework

    Anthropic says Claude Fable 5 has been redeployed globally and is now available to all users, alongside a fuller explanation of its cybersecurity safety classifiers. The company also introduced an early draft framework for grading AI jailbreak severity and opened a HackerOne program for researchers to report cyber jailbreaks.

  • 7The New Stack·7/1/2026security
    Anthropic restores Fable 5 access with tighter safety limits

    Anthropic will bring Fable 5 back to Claude products on July 1 after U.S. export controls were lifted, but access will initially be capped for subscription users and billed via usage credits for some enterprise tiers. The company says an improved safety classifier now blocks the bypass technique behind the incident in over 99% of cases, while keeping some cyberdefense capabilities available. Anthropic also said the underlying report did not reveal unique advanced cyber capabilities, and that similar vulnerabilities were found across other models tested, including Claude Opus 4.8, GPT-5.5, Kimi K2.7, and Claude Haiku 4.5.