Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size
Firstly a big thanks to the poster "hellohazine", he basically only removed the multi-lingual fat ofβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size
Firstly a big thanks to the poster "hellohazine", he basically only removed the multi-lingual fat ofβ¦
Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?
Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-serverβ¦
Is Microsoft-Phi dead?
Phi was one of my favorite models with a bit of a mixed reputation with some claiming it's benchmaxxβ¦
enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think
Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, β¦
Comparing 4bit quants for MLX
Curious what people think are the ideal 4-bit quantization types on MLX These quants seem to be the β¦
any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?
I have various quants of this model and am curious how they perform. can anyone recommend which bencβ¦
Building a budget 32GB β 48GB VRAM home AI server: 2-3x RX 9060 XT 16GB vs RTX 5060 Ti 16GB, AM5 vs used EPYC?
Iβm planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker serβ¦
I built a local realtime voice stack for Ollama: Parakeet STT β Qwen 2.5 7B β Qwen3-TTS
submitted by /u/InternationalGap3698 [link] [comments]β¦
Repeated generation is worth it and self-evaluation is effective
I made gemma4 12B write timestamp-anchored summaries of youtube video transcripts. I tested if the sβ¦
Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU
Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 β¦
Showoff Saturday: Local 4x 6000 Pro (multi-year progression)
Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama modelβ¦
My first run of Kimi K3 locally.
Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everytβ¦