Trendora

ScarfBench

Hold

Tools

An open benchmark for evaluating AI agents on enterprise Java framework migration tasks.

Why it's here

Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 7Hugging Face Blog·6/30/2026research
    ScarfBench Benchmarks AI Agents for Java Framework Migration

    Hugging Face Blog introduces ScarfBench, an open benchmark for evaluating AI agents on enterprise Java framework migration across Spring, Jakarta EE, and Quarkus. The benchmark measures whether migrated applications can build, deploy, and preserve behavior, and reports that current frontier agents still achieve less than 10% behavioral success.