Skip to content
r/LocalLLaMA · Communities

RTX 5080 + 32gb DDR4

Looking for some suggestions on what I should be realistically aiming for. Recently I've used qwen 3.6 27b at Q3KM, but I've also been able to run qwen3.6 35b at IQ2. 27B yields around 60 tk/s and 35B is yielding around 190tk/s. I know there is quite a bit of accuracy lost going down from Q6. I also have a spare 1080ti