Trendora

MBPP+

Assess

Tools

An improved programming benchmark for evaluating code-writing models.

Why it's here

Placed in Assess: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.

Evidence (1)

  • 8Hugging Face Blog·4/21/2026research
    QIMMA: a quality-first Arabic LLM leaderboard

    QIMMA is a new Arabic large language model leaderboard that validates benchmark quality before evaluating models. It consolidates 109 subsets from 14 benchmarks into a suite of more than 52,000 samples across seven domains, including cultural, STEM, legal, medical, safety, poetry, and coding tasks. The project also releases public per-sample outputs and includes Arabic-adapted code evaluation.