r/LocalLLaMA 🤖 Ai 👁 0

Cheapest setup for >10 tok/sec for 120B dense LLM

Hi all, I'm trying to wrap my head around hardware variables when it comes to LLM, and I have another question: what would be the cheapest way to run a 120B dense LLM at >10 tok/sec? I'm fine with Q5, ideally Q6 though.

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.