Skip to content
arXiv cs.LG · Papers

Maglev: Sliding Recurrent Memory

arXiv:2608.02870v2 Announce Type: replace Abstract: We introduce ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. ours{} consists of two coupled models: a prefiller $Q$, which leverages full attentionfootnote