proveKV – Honest 36× lossless (vs f32, 18x vs fp16) KV‑cache compression for LLMs (zero PPL regression)
I’m sharing a new open‑source repo that demonstrates a reproducible KV‑cache compression technique. - Result: 36× lossless / 68× lossy memory reduction vs. f32‑raw KV cache on SmolLM2‑1.7B + WikiText‑2 (0% ΔPPL). - Tr
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/LocalLLaMA
Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.