arXiv cs.CL
· Papers
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
arXiv:2607.02980v1 Announce Type: new Abstract: Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of dense attention. Chunk-wise sparse attention offers a promising alternative, but all existing methods fall short of full attention b