automated classifiers
AssessTechniques
Detection models used to flag potentially policy-violating content or abuse patterns.
Why it's here
Placed in Assess: 1 article(s) of evidence from 1 source(s), led by framework updates, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 6Anthropic News·4/24/2026framework_updateAnthropic updates election safeguards for Claude
Anthropic says it is strengthening Claude’s election safeguards ahead of the US midterms and other major elections in 2026. The company highlights bias evaluations, policy enforcement, automated detection, and testing that showed high compliance with election-related safety rules.