Trendora

SkillsBench

Hold

Tools

A community benchmark for measuring the effectiveness of AI agent skills.

Why it's here

Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 5The New Stack·7/6/2026research
    JetBrains finds Claude Code caveman mode saves far fewer tokens than claimed

    As AI coding tools increasingly use usage-based pricing, developers are trying to reduce token consumption by making assistant replies more terse. JetBrains tested a popular Claude Code skill that forces blunt, caveman-style output and found it reduced tokens by about 8.5% in real coding tasks, far below its claimed 65% savings because most agent output is still code, diffs, and exact tool output.