Skip to content
arXiv cs.LG · Papers

GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models

arXiv:2603.13418v2 Announce Type: replace Abstract: Structured pruning is widely applied to compress large language models (LLMs), but its performance depends heavily on how neuron importance is estimated. Most existing methods rely on activation statistics from a single calibration set, which introduces calibration bi