arXiv cs.LG
· Papers
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs
arXiv:2608.04048v1 Announce Type: new Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each target bit-width. We intr