Skip to content
r/LocalLLaMA · Communities

GLM-5.2 on 8xB200: the deployment math nobody spells out – NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside.

We have 8xB200 nodes and users keep asking us how to serve GLM-5.2 on them. Our engineering team went through everything published so far, and the optimal config is not the obvious one. Sharing the analysis because most of it applies wherever you rent or rack your B200s. The model GLM-5.2: ~750B total / ~40B active MoE