Skip to content
arXiv cs.CL · Papers

Less Is More: Reducing Token Counts Without Compromising Performance

arXiv:2506.15138v2 Announce Type: replace Abstract: Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost. Although longer, multi-word tokens can reduce fertility, naively adding them often degrades language model per