AMD Quark
HoldTools
A quantization toolkit used to compress model weights for faster inference and lower memory use.
Why it's here
Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 6Hacker News·7/3/2026researchWafer reports faster GLM5.2 inference on AMD MI355X
Wafer says it ran GLM5.2 on AMD MI355X with 2626 tokens per second per node at 2.4 RPS, while claiming the setup cost is more than 2x lower than Blackwell-based alternatives. The post also describes using MXFP4 quantization with AMD Quark and serving the model with sglang on ROCm after fixing speculative decoding support.