sycl : port multi-column MMVQ from CUDA backend (~45% speculative decoding speedup on Intel Arc) by masonmilby Β· Pull Request #21845 Β· ggml-org/llama.cpp
Saw this on other sub so posting here. For Intel ARC card holders. Big boost so update llama.cpp verβ¦