ARC-AGI
AssessPlatforms
A benchmark suite for evaluating AI systems on general intelligence and adaptive problem-solving.
Why it's here
Placed in Assess: 3 article(s) of evidence from 2 source(s), led by research-stage coverage, with 3 in the last 30 days. Confidence 42%.
Evidence (3)
- 8Hacker News·8/7/2026researchDeepSeek V4 Flash 0731 posts strong ARC-AGI benchmark results
DeepSeek announced V4 Flash 0731 with three reasoning variants and published benchmark results on ARC-AGI-1 and ARC-AGI-2. At maximum effort, the model reportedly scored 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2 with low per-task costs. The post includes detailed task-level pass/fail results and links to the paper and model page.
- 7OpenAI Blog·7/29/2026researchTwo settings tripled GPT-5.6 scores on ARC-AGI-3
OpenAI says two API settings significantly improved GPT-5.6 performance on the ARC-AGI-3 benchmark by preserving reasoning state and enabling compaction. The change increased scores and efficiency, showing how inference-time configuration can materially affect benchmark results.
- 6Hacker News·7/25/2026researchARC-AGI Leaderboard Tracks AI Reasoning Efficiency
The ARC-AGI leaderboard compares AI systems across ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3, with a focus on performance versus cost per task. It highlights how different model types, including reasoning systems, base LLMs, and Kaggle submissions, perform under varying compute constraints.