Skip to content
arXiv cs.CL · Papers

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

arXiv:2608.01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing t