Skip to content
arXiv cs.CL · Papers

ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabul