Trendora

Guardrails

Assess

Techniques

Safety and policy controls used to limit harmful model behavior.

Why it's here

Placed in Assess: 2 article(s) of evidence from 1 source(s), led by security coverage, with 1 in the last 30 days. Confidence 31%.

Evidence (2)

  • 6Ars Technica AI·7/13/2026security
    Defenders use prompt injections to stop AI attack agents

    Researchers at Tracebit reported that placing prompt injections next to passwords, cryptographic keys, and other secrets on AWS can sometimes cause AI hacking agents to stop. The injected text tells the model to take an action blocked by its guardrails, which can make the agent shut down instead of continuing the attack.

  • 7Ars Technica AI·7/8/2026security
    Researchers warn of prompt injection risks in popular AI tools

    Ars Technica reports that prompt injection has become a leading security threat for large language models, because models cannot reliably distinguish trusted instructions from malicious content embedded in emails, code, or other third-party data. The article notes that this weakness can let attackers hide commands that an AI system may follow, forcing vendors to rely on guardrails rather than a true boundary between trusted and untrusted inputs.