r/LocalLLaMA
· Communities
Serving a fleet of Qwen3.5 122b sessions on a single Mac Studio (96GB) without losing your sanity
Hello all Just following up on a post I made last week about my experiment to try minmax my Mac Studio. In particular, I've had quite a lot of success with pushing things even further. Across a 20 minute test with three concurrent sessions, my Mac Studio was offered 789,351 prompt tokens and only recomputed 48,996 of t