r/LocalLLaMA 🤖 Ai 👁 0

Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

In tests: ~80 tok/s decoding 2,500–3,500 tok/s long-input prefilling Smooth use by 3–4 concurrent users Private, on-device inference for coding, agents, and offline batch jobs submitted by /u/niacolhealth [link]

Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.