Skip to content
arXiv cs.CL · Papers

Accelerating Large Language Model Inference with Self-Supervised Early Exits

arXiv:2407.21082v3 Announce Type: replace Abstract: This paper presents a modular approach to accelerate inference in large language models (LLMs) by adding early exit heads at intermediate transformer layers. Each head is trained in a self-supervised manner to mimic the main model's predictions, allowing computation t