Gemma 4 12B is my new main squeeze
The Unsloth Q5_K_XL is officially my main squeeze for local coding. I started out with the Q4_K_XL, β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Gemma 4 12B is my new main squeeze
The Unsloth Q5_K_XL is officially my main squeeze for local coding. I started out with the Q4_K_XL, β¦
Can MTP models be used as standalone smaller models? (e.g. DS4 Flash/Pro)
I've been wondering about models that are trained with MTP (Multi-Token Prediction) and whether the β¦
PSA: You may not need to quantize spec draft when using MTP
Using `--spec-draft-type-k q4_0 --spec-draft-type-v q4_0` might actually decrease your context size!β¦
hello there! i made a tool to explore kokoro.
i built this on top of my own stack but the code is MIT for everything related. kokoro was pretty fuβ¦
Here is my llama.cpp NVFP4/MXFP6 GGUF quantizer tool
Hello everyone I wanted to share what I've been working on. I started writing NVFP4 kernels for llamβ¦
Magenta RealTime 2: Open & Local Live Music Models
Build and play AI musical instruments on your laptop! submitted by /u/phone_radio_tv [link] β¦
Finally finished my LLM server: EPYC 9575F, 4Γ RTX 3090 (96GB VRAM), 768GB ECC RAM
Took a while, but Nalthis is finally up and assembled. Specs: Supermicro H13SSL-N AMD EPYC 9575F (6β¦
Kimi K2.6 on 8ΓB200: expected vLLM/SGLang throughput?
Iβm planning to run moonshotai/Kimi-K2.6 on 8ΓNVIDIA B200 with vLLM or SGLang, likely using NVFP4(orβ¦
Qwen 3.6 35B on RTX 3080 10GB + 7700X + 32GB DDR5
Environment: GPU: RTX 3080 10GB CPU: Ryzen 7 7700x RAM: 32GB 6000mt/s OS: CachyOS engine: ik_llamacβ¦
Horus Image Generation is here! π€©π·
https://preview.redd.it/57kqog9iqd5h1.png?width=1537&format=png&auto=webp&s=85b3ec32b0797bdeb2a02108β¦
How are RTX 6000 PRO (Either WS/MaxQ/SE) prices going on your country/state?
Hello guys, hoping you're fine. I was wondering, how does the RTX 6000 PRO prices (in general for anβ¦
proveKV β Honest 36Γ lossless (vs f32, 18x vs fp16) KVβcache compression for LLMs (zero PPL regression)
Iβm sharing a new openβsource repo that demonstrates a reproducible KVβcache compression technique. β¦