Skip to content
r/LocalLLaMA · Communities

Strategies for capping thinking on ds4 flash 0731

I like the outputs from this model, but DAMN does it over think. Has anyone found a robust fix for this that isn't just capping output tokens? Anyone working on a 'thinking cap' for it? Some combo of llama params, or (system?) prompting technique? I'm all ears. submitted by /u/youcloudsofdoom [link] [comments]