Skip to content
arXiv cs.AI · Papers

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning

arXiv:2607.19358v1 Announce Type: new Abstract: Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. However, the O(n^2) computational complexity of standard self-attention causes inference costs to