arXiv cs.LG
· Papers
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
arXiv:2607.08786v1 Announce Type: new Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quality typically limits pruning to moderate un