MXFP4
AssessTechniques
A 4-bit floating-point quantization format used to reduce model size and improve inference efficiency.
Why it's here
Placed in Assess: 2 article(s) of evidence from 3 source(s), led by research-stage coverage, with 1 in the last 30 days. Confidence 47%.
Evidence (2)
- 8Simon Willison·7/27/2026model_releaseMoonshot releases Kimi K3 open weights, but deployment is limited
Moonshot AI has released the open weights for Kimi K3 on Hugging Face, making one of its largest language models available for self-hosting. The model targets long-horizon coding and knowledge work, but its 2.8-trillion-parameter MoE design and heavy hardware requirements mean only a small number of organizations can run it themselves.
- 6Hacker News·7/3/2026researchWafer reports faster GLM5.2 inference on AMD MI355X
Wafer says it ran GLM5.2 on AMD MI355X with 2626 tokens per second per node at 2.4 RPS, while claiming the setup cost is more than 2x lower than Blackwell-based alternatives. The post also describes using MXFP4 quantization with AMD Quark and serving the model with sglang on ROCm after fixing speculative decoding support.