KVarN: new KV-cache quant from Huawei. 3β5Γ KV cache compression with actual speed-up instead of slow-down, and unlike TurboQuant it holds up on reasoning (Apache 2.0, vLLM single flag)
The KV-cache quant race just got more interesting. Huawei just open-sourced KVarN, a KV-cache quantiβ¦