arXiv cs.AI
· Papers
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
arXiv:2603.15685v2 Announce Type: replace-cross Abstract: Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely on fixed window partitioning and attention-