r/LocalLLaMA 🤖 Ai 👁 0

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Requirements: M3 or new

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.