Trendora

Kimi K-series

Hold

Tools

Moonshot's large long-context mixture-of-experts language model family.

Why it's here

Placed in Hold: 1 article(s) of evidence from 2 source(s), led by research-stage coverage, with 1 in the last 30 days. Confidence 32%.

Evidence (1)

  • 7Hacker News·8/3/2026research
    Cloudflare scales Kimi and GLM with KV cache quantization and weight compression

    Cloudflare says it is serving Moonshot's Kimi K-series and Z.ai's GLM more efficiently on Workers AI by combining KV cache quantization, weight compression, and cache protection techniques. The company reports that FP8 KV cache and INT4 weights cut GPU memory use and lower costs while preserving model accuracy, using SGLang as the serving framework.