Skip to content
arXiv cs.LG · Papers

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs

arXiv:2608.04048v1 Announce Type: new Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each target bit-width. We intr