arXiv cs.LG
· Papers
Maglev: Sliding Recurrent Memory
arXiv:2608.02870v2 Announce Type: replace Abstract: We introduce ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. ours{} consists of two coupled models: a prefiller $Q$, which leverages full attentionfootnote