r/LocalLLaMA 🤖 Ai 👁 0

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B total, ~12B active, 3.4

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.