Gemma 4 QAT accuracy inconsistencies
Table from https://unsloth.ai/docs/models/gemma-4/qat#qat-analysis I heard that MoE models are usualβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Gemma 4 QAT accuracy inconsistencies
Table from https://unsloth.ai/docs/models/gemma-4/qat#qat-analysis I heard that MoE models are usualβ¦
Meme: every 'small local model after i open task manager
i swear every model is 'lightweight' until my laptop starts preparng for takeoff lol! submitted β¦
Does anyone know what PCIe mode was used for these benchmarks?
https://github.com/noonghunna/club-3090/blob/master/docs/DUAL_CARD.md It says PCIe only, but it doesβ¦
Better VRAM Estimator
This was for 32k context on both sides. I think the website was linked in the Wiki and the other is β¦
Cohere's unreleased coding model (early access for localllama)
Hey, Nick here from Cohere. Thanks for all the feedback on Command A+ the other week everyone. I reaβ¦
MoQ GGUFs and GSQ: Low-Bit GGUFs Are About to Get Much Better
submitted by /u/beneath_steel_sky [link] [comments]β¦
Experimentation with Qwen 3.6 and Gemma 4 - Guidance needed
Iβm a web developer doing mostly coding, but also project management, requirements analysis, testingβ¦
StepFun 3.7 Flash MTP Bench Strix Halo
This is the StepFun Step-3.7-Flash UD-IQ4_XS main model with the official StepFun MTP Q8_0 draft modβ¦
Dual GPUs - 3060 & 3090 on a P520
I've got a line on a reasonably priced 3090FE and I'm wondering whether it would play nicely with thβ¦
Gemma 4 QAT Q4_0 Bench on Strix Halo
Gemma 4 QAT Q4_0 Bench on Strix Halo These are Google's official Gemma 4 QAT Q4_0 GGUF models, serveβ¦
People from r/antiai must be barbaric
Just visited their sub, they are citing a 3 year old model answering without any tools about a can oβ¦
Activating MTP for QATGemma4 31b q4_0?
Has anyone figured out how to activate MTP for Gemma4βs new QAT q4_0 GGUF for 31b? Or is this still β¦