Run (your largest) local models from your iPhone
submitted by /u/BustyMeow [link] [comments]β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Run (your largest) local models from your iPhone
submitted by /u/BustyMeow [link] [comments]β¦
Qwen 3.6 27B 30GB Same top p: 98.358 Β± 0.033 % vs UD Q8 K XL 33GB Same top p: 97.426 Β± 0.041 %
This is not a diss to Unsloth, they make great quants and really move this community forward. I've bβ¦
Nemotron 3 Ultra. 550 billion parameters, 55B active. 1 million context
submitted by /u/AnticitizenPrime [link] [comments]β¦
I can fit 28% more context after building llama.cpp with OpenBLAS. Huh?
I've noticed a weird difference when building llama.cpp with the Vulkan and OpenBLAS backends vs. buβ¦
I accidentally crippled my 4x RTX 3090 LLM rig with a hidden PCIe 2.0 x4 slot and fixing it doubled Mistral 128B performance
Iβm posting this as a warning for anyone building multi-GPU local LLM rigs with older workstation/HEβ¦
NVIDIA Nemotron 3 Ultra is out.
Not sure how much this is in the "local" world but interesting what they are putting out. https://dβ¦
The DeepSWE benchmark was runned rather incompetently and the results are completely invalid
submitted by /u/Charuru [link] [comments]β¦
Unsloth on Apple Silicon- Pre-announcement announcement
submitted by /u/openSourcerer9000 [link] [comments]β¦
Nvidia's been paying shills on LinkedIn
3 different accounts, some even with LinkedIn Gold, made the above posts all on the same day. And clβ¦
Today made me realize just how bad things have gotten without Meta
submitted by /u/ForsookComparison [link] [comments]β¦
Qwen3.6-27B on 2x3090s: llama.cpp vs vLLM, all the flags, and the MTP acceptance/inference speed/context
written 20%-ish by me and 80% by Claude code Spent basically a whole day getting my box to run Qwen3β¦
VibeOS - Fully Hallucinated Operating System
Who needs programming anyway? submitted by /u/WhatererBlah555 [link] [comments]β¦