Trendora

IT-Bench

Hold

Tools

A benchmark for evaluating enterprise IT agent workflows such as SRE, security, and FinOps tasks.

Why it's here

Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 7Hugging Face Blog·2/18/2026research
    IBM and UC Berkeley Use IT-Bench and MAST to Explain Enterprise Agent Failures

    IBM Research and UC Berkeley studied failure patterns in enterprise agentic LLM systems on IT automation tasks using ITBench and the MAST failure taxonomy. Their analysis of 310 SRE traces found that incorrect verification was the most common failure driver, while some models failed cleanly and others suffered cascading errors or termination issues.