Kimi Delta Attention
AssessTechniques
An attention architecture used in Kimi models to improve information flow across long sequences.
Why it's here
Placed in Assess: 3 article(s) of evidence from 3 source(s), led by model releases, with 3 in the last 30 days. Confidence 49%.
Evidence (3)
- 4Hacker News·7/28/2026researchWalkthrough of the DeltaNet family of linear attention variants
This article explains how the DeltaNet family of linear attention methods is derived, starting from standard causal softmax attention and progressively simplifying it into linear attention forms. It then connects these ideas to DeltaNet, Gated DeltaNet, and Kimi Delta Attention, with an emphasis on the underlying state updates and how they can be implemented recurrently and chunkwise.
- 8Hacker News·7/28/2026model_releaseKimi K3 Architecture Highlights LatentMoE and NoPE
Kimi K3 is presented as a scaled-up production version of Kimi Linear, growing from 48B to 2.8T parameters and positioning itself as the largest open-weight model to date. The architecture emphasizes inference efficiency with components such as LatentMoE, Kimi Delta Attention, attention residuals, and a full switch to NoPE, while also adding native multimodal support.
- 8Simon Willison·7/16/2026model_releaseKimi K3 launches as an open 3T-class model
Kimi has introduced Kimi K3, a 2.8-trillion-parameter open model with native vision support and a 1-million-token context window. The company says it is optimized for long-horizon coding, reasoning, and knowledge work, with full weights planned for release on July 27, 2026.