r/LocalLLaMA 🤖 Ai 👁 0

LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU

As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine built from scratch. The

LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.