Skip to content
arXiv cs.LG · Papers

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eig