Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max β speed vs context depth, 100 turns, one graph
Setup: MacBook Pro M5 Max, 128 GB unified, macOS 26.5.2 Β· llama.cpp b10686 (Metal, 12 threads, batchβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max β speed vs context depth, 100 turns, one graph
Setup: MacBook Pro M5 Max, 128 GB unified, macOS 26.5.2 Β· llama.cpp b10686 (Metal, 12 threads, batchβ¦
What would you do with $4,000?
I already have a 5090 that I use got Hermes and coding mostly. My only jealously is trying models thβ¦
Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one
TL;DR: Every public low-bit GGUF of this model is secretly ~4.70 bpw. Shim the rows to 256 and it bβ¦
Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup
I have a 2 DGX Spark setup recently and I have been happily running Deepseek V4 Flash 0731. Since thβ¦
Any current Voice2Voice AI model that runs locally thatβs good?
(I mean STS) You guys remember sesame AI? With their really good AI voice model? Obviously ChatGPT hβ¦
67-84 t/s DeepSeek flash v4 off 2x GX10s
Finally achieved usable results with 2 gx10 at over 65 tokens a second sustained. The 2570 prompt evβ¦
llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference
Folks! We're just 50 PRs away from more faster inference. Hopefully by end of year. Experts!, pleaseβ¦
This finance-model benchmark card is more useful for what it discloses than for who "wins"
The official benchmark card for Ling-3.0-flash-Fin is a useful reminder that the unit being tested iβ¦
Tencent compressed Hy4-preview from 1.5TB to about 200GB GGUF and kept about 98% performance.
submitted by /u/RedditUsr2 [link] [comments]β¦
Someone tested various Models on the Political Compass test...
submitted by /u/Thrumpwart [link] [comments]β¦
Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering
Exo labs making some very exciting and interesting claims. The headline is bandwidth scales linearlyβ¦
Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)
I wanted to share my successful setup for running a Qwen 3.8 27B model with a massive context windowβ¦