Skip to content
HF Daily Papers · Papers

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering no flexibility to trade off quality for speed at inference time. Recent audio tokenize