r/LocalLLaMA 🤖 Ai 👁 0

Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?

Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the DeepSeek-V4-Flash

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.