Galaxy Z Fold6 as a local inference node β llama.cpp/Vulkan, homelab telemetry, SHA-256 model verification
Built a small Android app called Pocket Node that runs llama.cpp inference on-device. Here's what itβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Galaxy Z Fold6 as a local inference node β llama.cpp/Vulkan, homelab telemetry, SHA-256 model verification
Built a small Android app called Pocket Node that runs llama.cpp inference on-device. Here's what itβ¦
how to run gemma-4-12b-it-qat-w4a16-ct in vllm or any version quantized of the model
when running by using transformers it runs by using vllm some weird error come up plese can any bodyβ¦
What's your experience with Gemma4 QAT?
Hey everyone! Not a native speaker, so please correct my english where I make mistakes, (can only leβ¦
Is Gemma 4 12b good for coding?
How are you using it? Quantized? At what quantization level? On what hardware? Thank you for the infβ¦
Fully Unserious Post - Fully Hallucinated Operating System
Watched this with a mixture of disbelief and "this is brilliant/ludicrous" - if you're into offbeat β¦
Hear Me Out, Pi Fans Lurking Here
Not For Thee Maybe After watching several interviews with Pi's creator, Mario Zechner, I've come to β¦
club-3090 adds experimental FP8 support for Qwen3.6-27B!
Itβs finally here! Something many of us running dual RTX 3090 rigs have been anticipating. club-3090β¦
Preferred two LLM combo
Iβm using my MacBook Pro M1 Pro with 32GB to run Qwen3.5-35B in Q4 as my coding agent. I have a gamβ¦
llama-server router: a model pinned to one GPU still grabs a CUDA context on every card, so it OOMs when my others are full. Am I missing a flag or is this just how it is?
Running into something annoying with llama-server in router mode (`--models-preset`) and I can't telβ¦
I built a PyTorch MoE/MoD training framework with custom CUDA kernels [Apache 2.0]
PyTorch framework for training transformer LLMs with MoE and MoD architecture support, custom CUDA kβ¦
Qwen 3.6 27B on DeepSWE
Overview: It scored 2% (1.79% rounded up) It is 18/20th place scoring above Haiku 4.5 and Minimax Mβ¦
2-bit QAT model releases
So far model releases that take advantage of Quantization a Aware Training (QAT) have been focused oβ¦