Trendora

OSWorld

Assess

Tools

An evaluation benchmark for AI computer-use tasks in real software environments.

Why it's here

Placed in Assess: 4 article(s) of evidence from 5 source(s), led by model releases, with 1 in the last 30 days. Confidence 74%.

Evidence (4)

  • 8Anthropic News·7/24/2026model_release
    Anthropic Launches Claude Opus 5

    Anthropic introduced Claude Opus 5, its new default model for Claude Max and strongest model for Claude Pro. The company says it delivers major gains in coding, knowledge work, automation, and scientific tasks while improving cost-efficiency over Opus 4.8, though it still trails Mythos 5 on cybersecurity evaluations.

  • 7Anthropic News·2/25/2026funding
    Anthropic acquires Vercept to improve Claude computer use

    Anthropic has acquired Vercept, a team focused on perception and interaction systems for AI agents that operate inside live applications. The company says the deal will help advance Claude’s computer use capabilities, following recent gains in OSWorld performance with Claude Sonnet 4.6.

  • 8Anthropic News·2/17/2026model_release
    Anthropic Releases Claude Sonnet 4.6

    Anthropic has introduced Claude Sonnet 4.6, its most capable Sonnet model to date, with upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. The model adds a beta 1M-token context window, becomes the default on free and Pro plans, and keeps pricing aligned with Sonnet 4.5 while showing stronger resistance to prompt injection attacks.

  • 7Hugging Face Blog·2/3/2026model_release
    H Company Holo2-235B-A22B Preview sets new UI localization SOTA

    H Company released Holo2-235B-A22B Preview, its largest UI localization model to date, and reports new state-of-the-art results on ScreenSpot-Pro and OSWorld G. The company says agentic localization improves accuracy through iterative refinement, and that SkyPilot was used to coordinate large-scale training across multiple cloud providers.