Skip to content
r/LocalLLaMA · Communities

Don’t ignore llama.cpp RPC with old hardware. Results of a 5070 Ti and 1080 Ti over gigabit ethernet: it’s actually functional.

Results up front: I had to prioritize prefill or token generation - there was no happy medium. Using UD-Q4_K_XL, q8 kv cache, and 96k max context: focus on generation (MTP = 2): 350 pp and 36 tg @ 12k context focus on prefill (disabled MTP): 560 pp and 19 tg @ 12k context Focus on quality: using UD-Q5_K_XL, full kv cac