Trendora

Kimi Delta Attention

Assess

Techniques

An attention architecture used in Kimi models to improve information flow across long sequences.

Why it's here

Placed in Assess: 3 article(s) of evidence from 3 source(s), led by model releases, with 3 in the last 30 days. Confidence 49%.

Evidence (3)

  • 4Hacker News·7/28/2026research
    Walkthrough of the DeltaNet family of linear attention variants

    This article explains how the DeltaNet family of linear attention methods is derived, starting from standard causal softmax attention and progressively simplifying it into linear attention forms. It then connects these ideas to DeltaNet, Gated DeltaNet, and Kimi Delta Attention, with an emphasis on the underlying state updates and how they can be implemented recurrently and chunkwise.

  • 8Hacker News·7/28/2026model_release
    Kimi K3 Architecture Highlights LatentMoE and NoPE

    Kimi K3 is presented as a scaled-up production version of Kimi Linear, growing from 48B to 2.8T parameters and positioning itself as the largest open-weight model to date. The architecture emphasizes inference efficiency with components such as LatentMoE, Kimi Delta Attention, attention residuals, and a full switch to NoPE, while also adding native multimodal support.

  • 8Simon Willison·7/16/2026model_release
    Kimi K3 launches as an open 3T-class model

    Kimi has introduced Kimi K3, a 2.8-trillion-parameter open model with native vision support and a 1-million-token context window. The company says it is optimized for long-horizon coding, reasoning, and knowledge work, with full weights planned for release on July 27, 2026.