Skip to content
r/LocalLLaMA · Communities

Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar?

I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and was pretty damn happy with the results at first. But, Q8 quantization 130k context was barely squeezing through. Naturally, I bought two RTX 6000 Pros, just waiting for the next-gen releas