r/LocalLLaMA 🤖 Ai 👁 0

PSA: You may not need to quantize spec draft when using MTP

Using `--spec-draft-type-k q4_0 --spec-draft-type-v q4_0` might actually decrease your context size! With quantized spec draft, my context size is 83200. Without it (i.e. using the default of fp16 spec draft), context si

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.