r/LocalLLaMA 🤖 Ai 👁 0

Kimi K2.6 on 8×B200: expected vLLM/SGLang throughput?

I’m planning to run moonshotai/Kimi-K2.6 on 8×NVIDIA B200 with vLLM or SGLang, likely using NVFP4(or original QAT model). What real throughput should I expect for: Input length 8192 Ouput about 2048 concurrency 32 I’m

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.