arXiv cs.CL
· Papers
Less Is More: Reducing Token Counts Without Compromising Performance
arXiv:2506.15138v2 Announce Type: replace Abstract: Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost. Although longer, multi-word tokens can reduce fertility, naively adding them often degrades language model per