how to run gemma-4-12b-it-qat-w4a16-ct in vllm or any version quantized of the model
when running by using transformers it runs by using vllm some weird error come up plese can any body share the command of running it on vllm ? submitted by /u/SavingsWeather1659 [link] [comments]
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/LocalLLaMA
Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.