CyberGym
AssessTools
A benchmark for testing whether LLMs can reproduce known security vulnerabilities.
Why it's here
Placed in Assess: 1 article(s) of evidence from 1 source(s), led by security coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 8Anthropic News·3/6/2026securityAnthropic and Mozilla collaborate on Firefox security fixes
Anthropic says Claude Opus 4.6 discovered 22 Firefox vulnerabilities over two weeks, including 14 high-severity issues assigned by Mozilla. Mozilla used the reports and fixes to ship protections in Firefox 148.0, highlighting AI-assisted security research at accelerated speed.