r/LocalLLaMA 🤖 Ai 👁 1

Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 \ --n-cpu-moe 36 \ --no-mmap \ -ctk q8_0 \ -ctv q8_0 \ -fa 1 \ -c 9000 \ So rebuilt my llama.cpp build to run r

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.