r/LocalLLaMA 🤖 Ai 👁 0

Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s

Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k

Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.