Does CPU matter for GPU inference?
Hi, I'm currently building a PC which is exclusively going to be used for LLM inference. I'd like toβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Does CPU matter for GPU inference?
Hi, I'm currently building a PC which is exclusively going to be used for LLM inference. I'd like toβ¦
Qwen 3.6 for coding with 5090 - Your settings recommendations?
Hi, totally new to using LLMs for coding purposes, I am on Ubuntu and currently using LM Studio withβ¦
silx-ai/Quasar-Preview β’ Huggingface (5M context length)
https://huggingface.co/silx-ai/Quasar-Preview https://preview.redd.it/ur27udpzy66h1.jpg?width=900&foβ¦
Anyone seen benchmarks comparing Gemma 4 4-bit QAT vs. 8-bit standard quants?
I'm trying to find out if anyone has done any benchmarking comparing the Gemma 4 4-bit QAT models (vβ¦
Gemma 4 26B A4B IT QAT Comparison
Hopefully this isn't too low effort of a post. I just finished the benchmarks and I figured I'd postβ¦
ggml-webgpu: Improve prefill speeds for k-quants + refactor matmul for Q4/Q5/Q8 and k-quants by yomaytk Β· Pull Request #24225 Β· ggml-org/llama.cpp
This PR improves matmul performance for k-quants. The following table shows the improvement on the pβ¦
2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.
Forgive the claude summary, in the readme, but the base works. I'm still working on the hip kernal aβ¦
Jetbrains Mellum 2: a really good and performant model
Oh Hey Folks, I took the Mellum 2 model for a spin, so I wanted to share my impressions here. Disclβ¦
Sam Altman's Eye-Scanning Startup Lays Off Staff
submitted by /u/HumanDrone8721 [link] [comments]β¦
I fine-tuned Parakeet 0.6B for medical ASR β open weights, local Mac/CUDA/CPU
I fine-tuned NVIDIA's Parakeet TDT 0.6B v2 for clinical speech and am releasing the weights as Omi Mβ¦
Here's a llama.cpp CLI Command builder.
No accounts or sign up. No email requirements. No pop-ups and no cookies. No ads. Info is saved locaβ¦
New MLX LM Server From Apple
Key Technical Advantages: Performance: The M5 chip's neural accelerators significantly boost promptβ¦