r/LocalLLaMA
· Communities
How can i limit reasoning effort on the qwen3.5 and gemma4 models?
So, umhh, I am working on an agentic coding platform, and I need to make qwen3.5 and gemma4 models out of controlled reasoning chains. For example, at low, the model should prioritize finding the quickest solution and prioritize speed, and at high and xhigh, the model should absolutely hit that limit it has and figure