r/LocalLLaMA
· Communities
Strategies for capping thinking on ds4 flash 0731
I like the outputs from this model, but DAMN does it over think. Has anyone found a robust fix for this that isn't just capping output tokens? Anyone working on a 'thinking cap' for it? Some combo of llama params, or (system?) prompting technique? I'm all ears. submitted by /u/youcloudsofdoom [link] [comments]