New local AFM model is 20B
https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models 1β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
New local AFM model is 20B
https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models 1β¦
Pipeline parallelism in llama.cpp may be wasting your VRAM
By default, llama.cpp enables pipeline parallelism, presumably to speed up inference. In my testing,β¦
What harness are you guys using and for what use case?
Having a chat bot you can ask questions is cool and all but for more advanced stuff like tool callinβ¦
Is opencode subagents actually useful?
Does someone have a very good and simple opencode subagent setup or a tutorial video that would helpβ¦
Quick note on the QAT of recent
tldr: Googles quant is broken, use unsloth UD Q4_K_XL for now This might be low quality post, but ohβ¦
What is your best coding model on a DGX Spark?
My current setup > pi > unsloth/Qwen3.6-35B-A3B-GGUF > llamacpp what i got > ~50 tok/s > get most oβ¦
mtp: support for gemma-4 E2B and E4B assistants by max-krasnyansky Β· Pull Request #24282 Β· ggml-org/llama.cpp
MTP for tiny gemmas for mobiles or potatoes or raspberry Pi, or maybe for ants submitted by /uβ¦
GLM-5.1 and Kimi K2.6 THE CHEAPEST WAY TO RUN
Guys how to run it as cheap as possible to get at least 15-20 ts? Asking for a friend! As example 50β¦
16B dense on 16GB GPU vs 32B dense on 2x 16GB GPU
I'm currently trying to plan a build to run big(-ish) LLMs locally, and was wondering the following:β¦
Me: Arguing with an AI bot who just posted something on this sub about Llama 3.1.
For real tho, these bots need to turn on their web search functions and quit living in the past. Itββ¦
Qwen3.6-35B-A3B tool calling benchmark: ByteShape vs. Unsloth GGUFs, KV cache quants & long context performance
I've previously posted some small performance benchmarks, but this time I got interested in the qualβ¦
Was BitNet a dead end? What happened to ternary LLMs?
They seemed so promising at one point but the biggest ternary model is still 2B. What happened? Why β¦