TrajFM
HoldTechniques
A trajectory-level failure analysis pipeline that clusters and explains agent breakdowns.
Why it's here
Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 7Hugging Face Blog·1/21/2026researchAssetOpsBench benchmark targets industrial AI agent realism
Hugging Face Blog introduces AssetOpsBench, a benchmark for evaluating AI agents in industrial asset operations such as chillers and air handling units. It emphasizes multi-agent coordination, six qualitative scoring dimensions, and trajectory-level failure analysis to better reflect real-world operational constraints.