Kimi K2.6 on 8×B200: expected vLLM/SGLang throughput?
I’m planning to run moonshotai/Kimi-K2.6 on 8×NVIDIA B200 with vLLM or SGLang, likely using NVFP4(or original QAT model). What real throughput should I expect for: Input length 8192 Ouput about 2048 concurrency 32 I’m
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/LocalLLaMA
Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.