r/LocalLLaMA 🤖 Ai 👁 0

The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).

On the CPU, batch 1 is memory bandwidth bound. But if token/s = bandwidth / (bytes_per_weight * active_weights_per_token) the total number of parameters doesnt slow down the generation speed. So building the architecture

The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.