TAU-Bench
AssessTechniques
A benchmark referenced as an existing test for AI agents.
Why it's here
Placed in Assess: 1 article(s) of evidence from 1 source(s), led by framework updates, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 7The New Stack·7/9/2026framework_updateDevRev releases an enterprise AI agent benchmark
DevRev introduced an open Enterprise AI Agent Benchmark to better measure how AI agents perform on real enterprise work, especially across large, permissioned data contexts. The first release covers only the L1 and L2 tiers and includes the dataset, evaluation harness, judging criteria, results, and raw traces so others can reproduce the tests.