DeepSWE
AssessTools
A benchmark for challenging software engineering problem solving.
Why it's here
Placed in Assess: 3 article(s) of evidence from 2 source(s), led by model releases, with 3 in the last 30 days. Confidence 43%.
Evidence (3)
- 7Hacker News·8/12/2026model_releaseSpaceXAI releases Grok 4.6 with longer-task coding gains
SpaceXAI released Grok 4.6 less than a month after Grok 4.5, highlighting training that emphasizes self-checking, error correction and sustained progress on longer coding and agent tasks. Benchmark results show clear gains over Grok 4.5, but Grok 4.6 still trails the top models on several coding and terminal tests.
- 7The New Stack·8/5/2026researchMeta uses engineer code fixes to train internal AI coding tools
Meta is asking engineers to submit code fixes made while using its internal coding agent, MetaCode, so those corrections can help train and post-train future models. The effort has already produced more than 800 fixes from 7,000 weekly active users, and Meta says the data has improved Muse Spark 1.1 and will help train an upcoming model codenamed Watermelon.
- 7Hacker News·7/21/2026model_releasePoolside releases Laguna S 2.1 coding model
Poolside has released Laguna S 2.1, a 118B-parameter Mixture-of-Experts coding model with 8B active parameters per token and support for up to a 1M-token context window. The company says it is optimized for longer-horizon agentic coding work and reports strong results on benchmarks such as Terminal-Bench 2.1, SWE-Bench Multilingual, SWE-Bench Pro, DeepSWE, SWE Atlas, and Toolathlon Verified.