Trendora

SlopCodeBench

Assess

Tools

A benchmark for measuring coding-agent behavior and code quality.

Why it's here

Placed in Assess: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 1 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 6Hacker News·7/27/2026research
    Benchmarking Opus 5 on SlopCodeBench

    This post reports benchmark results for Opus 5 on SlopCodeBench, a coding-focused evaluation. It discusses how the model performs on agentic programming tasks and what the benchmark reveals about coding quality and failure modes.