llama.cpp: Hy3 PR + GGUFs
Early stages as the model was just released yesterday, but seems to be working already. Yay! Getting coherent output from the Q2_K, at about 10-11t/s on a 5090 + Zen 4 w/ 96GB DDR5. https://github.com/ggml-org/llama.cpp/
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/LocalLLaMA
Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.