r/LocalLLaMA 🤖 Ai 👁 0

how to run gemma-4-12b-it-qat-w4a16-ct in vllm or any version quantized of the model

when running by using transformers it runs by using vllm some weird error come up plese can any body share the command of running it on vllm ? submitted by /u/SavingsWeather1659 [link] [comments]

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.