Dumb question: How would performance be if you took a used server with like 80 lanes pcie 5 and stuck NVMe on them for model run?
So for LLMs, VRAM speed is king. But what if you bought a used server which had, for example, 80 lan…
AI tools, cybersecurity and development news aggregated from top sources — saved permanently with unique URLs.
Dumb question: How would performance be if you took a used server with like 80 lanes pcie 5 and stuck NVMe on them for model run?
So for LLMs, VRAM speed is king. But what if you bought a used server which had, for example, 80 lan…
Tried to benchmark Google’s new on-device dictation models (Eloquent) and basically couldn’t
I tried to benchmark Google’s new on-device dictation app (Eloquent) and basically couldn’t. It drop…
Looking for small rack or shelf for Sparks / Mac Studio / Halo Strix devices that host my llms.
I have two dgx sparks and a framework strix halo computer sprawled out over a wire shelf and am look…
Hot Take "Rigid code is better than Flexible code if you're on a budget"
I've spent the last six months trying to build a fully local, agentic pipeline for a text_processing…
Are these quants of QAT better than non-QAT? What do I use?
https://huggingface.co/mradermacher/gemma-4-31B-it-qat-q4_0-unquantized-i1-GGUF/tree/main https://hu…
Best Open-Source AI coding model for my specs?
hello everyone! im looking for the most powerful open-source coding ai while still fitting my system…
Remove padding and multiple D2D copies for MTP by gaugarg-nv · Pull Request #24086 · ggml-org/llama.cpp
Another day, another MTP speedup submitted by /u/jacek2023 [link] [comments]…
Local LLM good for OCR of handwriting?
I am using qwen3-vl:8b and ollama for doing OCR on scans of handwritten letters and it is doing a de…
DeepMind Just Dropped "DiffusionGemma" — Text Generation via Image-Style Diffusion Model
Another open weight model got dropped today, this one's from DeepMind, seems like a good day for the…
Model/tooling recommendations for complex document processing.
I have huge stacks of mill test reports for metal shipments. Each test report is 1-5 pages, in what …
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottlen…
Any recent news/updates on taalas chips?? They said they gonna bake the mid tier llm model into their chip.
They said in spring they're gonna bake or hardcode a mid tier Llm into their chip. Do anyone have an…