r/LocalLLaMA
· Communities
Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4, what hardware are you running?
Trying to see what I can scrounge together bare minimum hardware requirements to get up to that rough speed. Right now I'm running an RX6600XT and Ryzen 7 5700X with 32GB of DDR4 at 3600MHZ. CachyOS, vanilla llama.cpp built with ROCm and a workaround going to make it work with my GPU. Works about 30-35% faster on prefi