Skip to content
arXiv cs.CL · Papers

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

arXiv:2607.02980v1 Announce Type: new Abstract: Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of dense attention. Chunk-wise sparse attention offers a promising alternative, but all existing methods fall short of full attention b