Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.
In tests: ~80 tok/s decoding 2,500β3,500 tok/s long-input prefilling Smooth use by 3β4 concurrent usβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.
In tests: ~80 tok/s decoding 2,500β3,500 tok/s long-input prefilling Smooth use by 3β4 concurrent usβ¦
The Session You Cannot Take With You | EARENDIL
submitted by /u/MoneyPowerNexis [link] [comments]β¦
LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU
As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multiβ¦
DGX Spark now sells for 6000-8000 euros. I still remember when it was just 4000.
submitted by /u/Afraid-Yoghurt6731 [link] [comments]β¦
Given the MiniMax H3 LoRAs Debacle - Some Important Context for Censorship enforcement and laws in China
*I felt the need to write this post because it seems like very few people on this sub are aware of Cβ¦
MiniMax issues
https://www.reddit.com/r/StableDiffusion/s/HrU7odaJe6 I think this is more important that all the poβ¦
Qwen Developers' responses from their recent Twitter/X AMA
Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combinedβ¦
I updated my localy run benchmark with DeepSeek V4 Flash 0731
It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartβ¦
Qwen3-TTS voice cloning is now in mainline llama.cpp β the old demo finally became real support
People may remember the Qwen3-TTS llama.cpp demo from a few months ago. That PR said it probably wouβ¦
Chinaβs Open-Weight Models Will Be Spared US Safety Tests
submitted by /u/fallingdowndizzyvr [link] [comments]β¦
White House AI Guidelines Exempt U.S. Open Models From Government Review
submitted by /u/realmvp77 [link] [comments]β¦
Kimi K3 full model running on 16x GB10 cluster at 20+tps
Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tpβ¦