Skip to content
r/LocalLLaMA · Communities

GLM 5.2 FP8 with FP8 KV – Terminal-Bench 2.1 = 79.8 (with one time-out that I didnt re-run)

I wanted to test the official results vs fp8 + fp8 kv. basic sglang setup on H200. If anyone wants one of the official tests do ping me. I didnt rerun the one so it might go up a bit :) TERMINAL-BENCH 2.1 — FINAL RESULTS (mymodel via mini-swe-agent) TOTAL: 89 tasks PASSED: 71 (79.8%) FAILED: 17 ERRORED: 1 Input tokens