Dense vs MoE quantization resiliance
Which one is more resiliant to quantization? Especially at 4-bit? My experience:i tried gemma4 26b aβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Dense vs MoE quantization resiliance
Which one is more resiliant to quantization? Especially at 4-bit? My experience:i tried gemma4 26b aβ¦
Cool stuff to do with NVIDIA RTX 6000 PRO 96GB VRAM
I have been a C++ dev for 3 years as long as have done PyTorch in my free time (not that good in theβ¦
Alternatives to ChromaDB for easy RAG search
I'm disappointed that ChromaDB's local, free "single node" version is still getting second-class, haβ¦
GraphKV, kv cache optimization based on graph embedding models
I've been working on a project inspired by TurboQuant, It isnt perfect but it's pretty good for a prβ¦
5 Months Later: open-deepthink Now Has Full Knowledge Distillation Mode
Hey r/LocalLLaMA, Some of you might remember when I posted about this project back around September β¦
I can't wait for all the x250 sample distills of Mythos and GPT-5.6
Just kidding. Are there any distills that actually improve a model's quality? I remember the Qwen R1β¦
Gemma4 12B - Experiences?
Anyone check out the new Gemma4 12B that dropped 3 days ago? Integrated vision and audio recognitionβ¦
Gemma 4 31B QAT Q4 vs standard Q4 β Top1 KLD benchmark results have me confused. Someone please explain or poke holes in this.
I'll be upfront: I vibe-benched and vibe-reported this with Claude Sonnet 4.6, but I reviewed and edβ¦
dvlt.cu: inference engine written from scratch in CUDA/C++ for NVIDIA's DVLT 3D transformer model
Im into both HPC and 3D reconstruction, so I built this as a side project. dvlt.cu is a single 5MB bβ¦
Introduction to LLM API Benchy
As i was struggling to find a good benchmark for my LLM and inference engines and always did somethiβ¦
QAT MTP Heads Upload + PARALLEL=2 Fix + 12B 2-slot Bench
Title: Gemma 4 QAT MTP assistant heads now public on HuggingFace + PARALLEL=2 crash fix + 12B 2-slotβ¦
Are local models good enough to replace Claude/Codex solely for simple HTML tasks?
I know local models canβt compete fully yet, but Iβm curious about where the limits are. My use caseβ¦