Trendora

ARC-AGI

Assess

Platforms

A benchmark suite for evaluating AI systems on general intelligence and adaptive problem-solving.

Why it's here

Placed in Assess: 3 article(s) of evidence from 2 source(s), led by research-stage coverage, with 3 in the last 30 days. Confidence 42%.

Evidence (3)

  • 8Hacker News·8/7/2026research
    DeepSeek V4 Flash 0731 posts strong ARC-AGI benchmark results

    DeepSeek announced V4 Flash 0731 with three reasoning variants and published benchmark results on ARC-AGI-1 and ARC-AGI-2. At maximum effort, the model reportedly scored 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2 with low per-task costs. The post includes detailed task-level pass/fail results and links to the paper and model page.

  • 7OpenAI Blog·7/29/2026research
    Two settings tripled GPT-5.6 scores on ARC-AGI-3

    OpenAI says two API settings significantly improved GPT-5.6 performance on the ARC-AGI-3 benchmark by preserving reasoning state and enabling compaction. The change increased scores and efficiency, showing how inference-time configuration can materially affect benchmark results.

  • 6Hacker News·7/25/2026research
    ARC-AGI Leaderboard Tracks AI Reasoning Efficiency

    The ARC-AGI leaderboard compares AI systems across ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3, with a focus on performance versus cost per task. It highlights how different model types, including reasoning systems, base LLMs, and Kaggle submissions, perform under varying compute constraints.