SEL-BIRD
HoldTools
A tool collection with specialized and domain-expanded operations for benchmark tasks.
Why it's here
Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 6Hugging Face Blog·4/15/2026researchVAKRA benchmarks agent reasoning and tool use in enterprise workflows
Hugging Face Blog introduces VAKRA, an executable benchmark for evaluating AI agents in enterprise-like environments with API and document-based tasks. The post also outlines dataset details and observed failure modes, noting that models perform poorly on multi-step reasoning and tool-use workflows.