arXiv cs.CL
· Papers
Accelerating Large Language Model Inference with Self-Supervised Early Exits
arXiv:2407.21082v3 Announce Type: replace Abstract: This paper presents a modular approach to accelerate inference in large language models (LLMs) by adding early exit heads at intermediate transformer layers. Each head is trained in a self-supervised manner to mimic the main model's predictions, allowing computation t