r/LocalLLaMA
· Communities
Decrease the power limit of your 5090 to at least 480W – the performance penalty for inference is negligible.
I run my inference machine in the living room, so noise and heat output are a significant concern. Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% less t/s in decode and 8.8% in prefill (which is already very fast). Well worth the massive noise reduction, heat output