arXiv cs.LG
· Papers
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts
arXiv:2607.20519v1 Announce Type: new Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block. Learned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time weighting of pe