Serving TTS/cloning models on llama.cpp?
Are there any quality voice cloning and speech generation models that already have support in Llama.β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Serving TTS/cloning models on llama.cpp?
Are there any quality voice cloning and speech generation models that already have support in Llama.β¦
A cooling chamber for dgx spark and gb10 machines at computex 2026
submitted by /u/rexyuan [link] [comments]β¦
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
Up to 5.8x throughput speedup on Qwen3 Paper : https://arxiv.org/abs/2605.29707 Code : https://githβ¦
Has there been any recent new development on which quant is considered optimal?
I recall in earlier days, q4 was said to be optimal. That is to say, if you have a: small q8 modelβ¦
Iβm upsetβ¦
So long story short - openai 20$ subscription is much better than my local AI stackβ¦ r7900xtx+32GB Rβ¦
Qwen 3.6 27B MTP - Adding spec-type and spec-draft-n-max is dropping tps and reducing GPU utilization
I have a 5090 power limited to 475W. When I run the following command, it barely hits 300W and I getβ¦
Big week for open AI, with 25+ notable open-weight drops across every modality (from Victor M on π)
π§ LLMs β NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFβ¦
JSON string errors caused by 4-bit quant or KV Cachce quant?
500 Failed to parse tool call arguments as JSON: [json.exception.parse_error.101] parse error at linβ¦
Local vs Frontier on low-level systems engineering
Hey r/LocalLLaMA, Before anyone jumps on me, this is absolutely not a post about how great Qwen is πβ¦
Qwen3.6-35B-A3B-Uncensored-Claude-4.6-Genesis-APEX-GGUF
Here model: https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Claude-4.6-Genesis-APEX-GGβ¦
DeepSeek V4 Flash is amazing! (WIP llama.cpp PR #24162)
In case you're not aware already, the DeepSeek V4 series is finally getting supported on llama.cpp wβ¦
AA comparison of the latest local models
I picked models I consider local (usable on 3Γ3090), so there are no 300B models, and you should proβ¦