Trendora

AI Safety Levels

Assess

Techniques

A tiered safety classification used to define required safeguards as model capabilities increase.

Why it's here

Placed in Assess: 1 article(s) of evidence from 1 source(s), led by framework updates, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 7Anthropic News·2/24/2026framework_update
    Anthropic updates Responsible Scaling Policy to version 3.0

    Anthropic has released version 3.0 of its Responsible Scaling Policy, the voluntary framework it uses to reduce catastrophic AI risks. The update aims to strengthen what has worked, fix gaps in the prior policy, and improve transparency and accountability in model-development decisions.