Skip to content
r/LocalLLaMA · Communities

GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone at 1-5 tok/s: +60GB model, 11GB of RAM, CPU only

This is a OnePlus 15R with about 11GB of usable RAM. The heaviest model is gpt-oss-120b, Q4_K_M, 60GB on disk. So it's roughly 5x bigger than the memory it's running in, which means keeping it resident isn't a matter of tuning, it just can't happen. It runs anyway: 1.3 tok/s at the model's own routing width (default ex