Skip to content
arXiv cs.CL · Papers

High-Layer Attention Pruning with Rescaling

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional training-free structured pruning methods often employ a heuristic metric that indiscriminately removes some attention h