I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM.
Hey guys, Just finished benchmarking DeepSeek V4 Flash 284B + DSpark on a single RTX PRO 6000 96GB. β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM.
Hey guys, Just finished benchmarking DeepSeek V4 Flash 284B + DSpark on a single RTX PRO 6000 96GB. β¦
Qwen3.6 35B (2 min) vs Muse Glimmer 30B (4 min) on custom Llama.cpp build (RTX 5080)
Muse Glimmer 30B feels significantly more precise and reliable, it almost never drops the ball or brβ¦
Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s
I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layeβ¦
Qwen 3.8 27B β MTP or DFlash?
Do we.know whether the 27B model will ship with a DFlash or MTP head? It's super exciting, but sinceβ¦
LFM2.5-VL-3B recognizes Steve from Minecraft running locally on an iPhone 17
Liquid AI put out LFM2.5-VL-3B today, which is a 3.1B vision model that weighs roughly 2GB and fits β¦
DeepSeek V4 Flash 0731 uncensored (jailbreak pt2)
Since lot's of people were sceptical or whatever, heres how to uncensor / jailbreak V4 flash and proβ¦
Meta's Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark
Been tinkering with speculative decoding on Apple Silicon for a while, and this week I got Meta's neβ¦
Today is Models Day
submitted by /u/Fz1zz [link] [comments]β¦
CohereLabs/North-Micro-Vision-Instruct Β· Hugging Face
North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resoluβ¦
Which Qwen3.8 model size do you want the most?
Just wanna get a sensing of the hardware ownership spread in the sub. I could ask that directly, butβ¦
Best models 14b and smaller as of today?
For the GPU impoverished submitted by /u/Thatisverytrue54321 [link] [comments]β¦
Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show
Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlaβ¦