Trendora

red teaming

Trial

Techniques

Adversarial testing used to uncover failures, vulnerabilities, and unsafe behaviors.

Why it's here

Placed in Trial: 5 article(s) of evidence from 6 source(s), led by regulatory news, with 3 in the last 30 days. Confidence 81%.

Evidence (5)

  • 9The New Stack·7/29/2026regulation
    Anthropic and peers urge slowdown on frontier AI development

    More than 1,100 AI researchers and executives signed an open letter asking governments to create a mechanism to intentionally slow frontier AI development if safety measures cannot keep pace. Anthropic publicly backed the petition, citing its research on recursive self-improvement, while OpenAI also expressed support and U.S. lawmakers considered emergency control legislation.

  • 3Simon Willison·7/25/2026research
    Boris Cherny on Opus 5’s resistance to prompt injection

    Boris Cherny said that Opus 5 is the least prompt-injectable model he has seen, based on PI evaluations and red-teaming. He noted this as a standout property beyond the model’s benchmark scores, and pointed readers to the relevant system card section.

  • 7OpenAI Blog·7/15/2026research
    GPT-Red: Automated Red Teaming for Safer AI

    OpenAI introduces GPT-Red, an automated red teaming system that uses self-play to find weaknesses in AI systems. The approach is designed to improve robustness against prompt injection while supporting broader safety and alignment efforts.

  • 9The New Stack·6/12/2026regulation
    Anthropic to disable Fable 5 and Mythos 5 after U.S. export-control directive

    Anthropic says the U.S. government ordered access to Fable 5 and Mythos 5 suspended for all foreign nationals, forcing the company to disable the models for all customers to remain compliant. Anthropic says the directive was based on a claimed jailbreak concern, but it believes the cited issue is a narrow vulnerability already observable in other publicly available models and not a universal bypass.

  • 6OpenAI Blog·4/23/2026security
    GPT-5.5 Bio Bug Bounty

    OpenAI is running the GPT-5.5 Bio Bug Bounty, a red-teaming challenge focused on finding universal jailbreaks that could bypass bio safety protections. The program offers rewards of up to $25,000 for valid findings.