Trendora

AI jailbreak severity framework

Assess

Techniques

A proposed framework for categorizing how severely a jailbreak can bypass an AI model's safeguards.

Why it's here

Placed in Assess: 1 article(s) of evidence from 1 source(s), led by framework updates, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 7Anthropic News·7/2/2026framework_update
    Anthropic details Fable 5 cyber safeguards and jailbreak severity framework

    Anthropic says Claude Fable 5 has been redeployed globally and is now available to all users, alongside a fuller explanation of its cybersecurity safety classifiers. The company also introduced an early draft framework for grading AI jailbreak severity and opened a HackerOne program for researchers to report cyber jailbreaks.