Skip to content
r/LocalLLaMA · Communities

Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can’t get rid of endless thinking

Title pretty much says it all. I’ve deployed Qwen 3.8 27B using vLLM on an RTX 6000 Pro (tried multiple vLLM releases and launch recipes), but I can't get it into a usable state because of crazy long reasoning passes. Regardless of the thinking effort setting (xhigh, medium, or low), it takes way too long to respond: x