arXiv cs.CL
· Papers
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
arXiv:2608.01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing t