arXiv cs.LG
· Papers
Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
arXiv:2607.18284v1 Announce Type: new Abstract: To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requirements restricting their applicability in resource-constrained environments. To address the problem of neural network (NN) compression S