Gemma 4 QAT + MTP: max 33% speed increase in token generation, any ideas?
Hello, My setup is 2x RTX 3060 Ti 8GB, without the assistant model (MTP) I get around 75t/s, adding β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Gemma 4 QAT + MTP: max 33% speed increase in token generation, any ideas?
Hello, My setup is 2x RTX 3060 Ti 8GB, without the assistant model (MTP) I get around 75t/s, adding β¦
Looking for a local "NotebookLM for lawyers" setup β what am I doing wrong?
Hello everyone I am totally new to LocalLLMs and only used chatGPT/Claude/NotebookLM before. So bearβ¦
OpenEnv is now owned by HF, Torch, Prime Intellect, Unsloth, Modal, Mercor, and more! Use it for training agents.
OpenEnv is a tool for creating an agentic execution environment like terminals, browsers, or anythinβ¦
[3090] Gemma4 QAT + MTP quick TPS numbers [TLDR 1.2-1.8x better]
These last few weeks have been godsend for 24GB (and below) gpu poor peeps. Killer models released β¦
llama-launcher Release
Hello everyone, I've been working on a point and click GUI to make tinkering with llama-server flagsβ¦
mtmd : add video input support by ngxson Β· Pull Request #24269 Β· ggml-org/llama.cpp
Show your videos to Gemma or Qwen today submitted by /u/jacek2023 [link] [comments]β¦
Gemma 4 Chat Template now has preserve thinking
submitted by /u/seamonn [link] [comments]β¦
Used local Ollama (gemma4:e4b + nomic-embed-text) to bulk-generate AI summaries for 4300 arXiv papers and push them to a remote Cloudflare DB β pipeline walkthrough
I built ArxivExplorer, a semantic arXiv search engine with AI-generated summaries. The live version β¦
whatβs was your local daily driver for coding last week?
drop your favorite model and quant in the comments. View Poll submitted by /u/be566 [link] β¦
kv-cache : avoid kv cells copies by ggerganov Β· Pull Request #24277 Β· ggml-org/llama.cpp
Improved MTP performance (For Gemma-4) This got merged yesterday. Available b9551 onwards. submitβ¦
Most reliable way to do PDF to JSON?
Hello everyone, I am currently stuck at automating a process where I need to parse medium-hard levelβ¦
[Benchmark] DFlash Speculative Decoding + KV Cache Compression on RTX 5090 β 3.26x Speedup
Hardware: RTX 5090 | Model: Qwen3.6-27B | Framework: BeeLlama.cpp Full benchmark scripts, raw data, β¦