The open-weights carousel never stops.
submitted by /u/InternationalGap3698 [link] [comments]β¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
The open-weights carousel never stops.
submitted by /u/InternationalGap3698 [link] [comments]β¦
Kimi K3 for local use (1.56TB β 594GB) compressed and released by Unsloth
The model was quantized to 8, 4, 2, and 1 bit. Characteristics: Q8: 8-bit 1.56 TB, lossless Q4: 4-bβ¦
I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base Β· Hugging Face
I know this is the 1000000th new sub billion parameter model out there and probably isn't as good asβ¦
PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
If your GGUF has MTP/NextN tensors baked in (GLM-5.2, hy_v3, qwen35moe, step35, etc.), recent llama.β¦
Ilintar's Official Guide To Model Selection
Inspired by multiple discussions here and on some Discords I frequent, I've decided to share with yoβ¦
Everyone posts day-one impressions. What's still in your stack a month later?
Day one threads are the least useful thing we produce here and we produce a lot of them. Model dropsβ¦
Has anyone tried Qwen3.7 flash on openrouter? How does it compare to our Qwen 3.6 27B?
This might be the next open weight release by qwen team. What you feel like is improved or have becoβ¦
First Kimi K3 results on home lab ~ 4t/s
I've got better results than expected for 768gb DDR5 and 2x5090. Using fork https://github.com/pwilβ¦
I keep coming back to Qwen... Over and Over. Is there really nothing better under 120B?
I was looking for a strong coding model and a strong general model, both should be 120b or under. Afβ¦
"Uncensored" LLMs are measurably more optimistic than their base models
Hi. Many people think uncensored models are basically the same model that just doesn't refuse, but..β¦
The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).
On the CPU, batch 1 is memory bandwidth bound. But if token/s = bandwidth / (bytes_per_weight * actiβ¦
Understand Kimi K3 from first principles: a recommended order for anyone trying to understand this beast
Everyone is talking about Kimi K3, but if you jump straight into the technical report, youβll quicklβ¦