red teaming
TrialTechniques
Adversarial testing used to uncover failures, vulnerabilities, and unsafe behaviors.
Why it's here
Placed in Trial: 5 article(s) of evidence from 6 source(s), led by regulatory news, with 3 in the last 30 days. Confidence 81%.
Evidence (5)
- 9The New Stack·7/29/2026regulationAnthropic and peers urge slowdown on frontier AI development
More than 1,100 AI researchers and executives signed an open letter asking governments to create a mechanism to intentionally slow frontier AI development if safety measures cannot keep pace. Anthropic publicly backed the petition, citing its research on recursive self-improvement, while OpenAI also expressed support and U.S. lawmakers considered emergency control legislation.
- 3Simon Willison·7/25/2026researchBoris Cherny on Opus 5’s resistance to prompt injection
Boris Cherny said that Opus 5 is the least prompt-injectable model he has seen, based on PI evaluations and red-teaming. He noted this as a standout property beyond the model’s benchmark scores, and pointed readers to the relevant system card section.
- 7OpenAI Blog·7/15/2026researchGPT-Red: Automated Red Teaming for Safer AI
OpenAI introduces GPT-Red, an automated red teaming system that uses self-play to find weaknesses in AI systems. The approach is designed to improve robustness against prompt injection while supporting broader safety and alignment efforts.
- 9The New Stack·6/12/2026regulationAnthropic to disable Fable 5 and Mythos 5 after U.S. export-control directive
Anthropic says the U.S. government ordered access to Fable 5 and Mythos 5 suspended for all foreign nationals, forcing the company to disable the models for all customers to remain compliant. Anthropic says the directive was based on a claimed jailbreak concern, but it believes the cited issue is a narrow vulnerability already observable in other publicly available models and not a universal bypass.
- 6OpenAI Blog·4/23/2026securityGPT-5.5 Bio Bug Bounty
OpenAI is running the GPT-5.5 Bio Bug Bounty, a red-teaming challenge focused on finding universal jailbreaks that could bypass bio safety protections. The program offers rewards of up to $25,000 for valid findings.