Trendora

Artificial Analysis

Trial

Platforms

A benchmarking and comparison source for model performance and pricing.

Why it's here

Placed in Trial: 6 article(s) of evidence from 2 source(s), led by model releases, with 5 in the last 30 days. Confidence 53%.

Evidence (6)

  • 8Hacker News·8/12/2026model_release
    Grok 4.6 reaches AI index frontier at lower cost

    Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, tying the GPT-5.6 Sol max score and narrowing the gap to the top frontier models. It also shows strong agentic performance across knowledge work, banking, and terminal benchmarks while keeping pricing at $2/$6 per million input/output tokens.

  • 6Hacker News·8/6/2026research
    Qwen3.8 Max tops Artificial Analysis's agentic index

    Artificial Analysis updated its Intelligence Index to v4.1.1 and published a new evaluation for Qwen3.8 Max. The model is now ranked as the best overall on the agentic index, based on the site's independent benchmark suite. The update also includes methodology changes and refreshed scores for other models, but the headline result is Qwen3.8 Max moving to the top of the ranking.

  • 5Hacker News·7/31/2026model_release
    DeepSeek V4 Flash 0731 ranks highly in model analysis

    Artificial Analysis reports that DeepSeek V4 Flash 0731 (Reasoning, Max Effort) scores strongly on its Intelligence Index, placing it among the leading open-weight models in its size class. The model has a 1M-token context window, 13B active parameters out of 284B total, and relatively low input/output pricing compared with peers.

  • 6Hacker News·7/24/2026research
    Opus 5 Tops Artificial Analysis Intelligence Leaderboard

    Artificial Analysis reports Claude Opus 5 as the highest-scoring model on its Intelligence leaderboard, with Opus 5 variants ranked ahead of other leading systems such as Claude Fable 5 and GPT-5.6 Sol. The page compares models across intelligence, speed, latency, price, and context window using its multi-evaluation methodology.

  • 4Hacker News·7/19/2026research
    How one researcher cut AI agent token costs with shared subscriptions

    A Quesma researcher describes building a deep-research pipeline for studying AI agent economics after an initial run exhausted a Claude Max plan in 30 minutes. The revised setup uses Claude Code as the main harness, with Codex and Antigravity added as headless subagents sharing memory through claude-mem, and cheaper models assigned to specific roles to reduce cost and improve verification.

  • 8Simon Willison·6/17/2026model_release
    GLM-5.2 open weights model lands with 1M context

    Z.ai has released GLM-5.2 as open weights under an MIT license after an earlier limited rollout to coding subscribers. The text-only Mixture-of-Experts model has a 1 million token context window and is being reported as the leading open weights model on independent benchmarks, though it is relatively token-hungry.