r/LocalLLaMA
· Communities
For RAG specifically, prefill speed matters more than decode and why Strix Halo struggles for interactive use
Seeing a lot of "what hardware for local RAG" threads lately, and the framing that keeps getting missed is: decode tok/s is not the bottleneck for RAG. Prefill is the bottleneck for RAG. RAG queries stuff thousands of tokens of retrieved context into every prompt. On unified memory boxes like Strix Halo, prefill throug