Skip to content
arXiv cs.CL · Papers

Hyperloop Transformers

arXiv:2604.21254v3 Announce Type: replace-cross Abstract: LLM architecture research generally aims to maximize model quality subject to fixed compute/latency budgets. However, many applications of interest such as edge and on-device deployment are further constrained by the model's memory footprint, thus motivating par