Skip to content
arXiv cs.AI · Papers

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

arXiv:2603.18558v2 Announce Type: replace-cross Abstract: Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows. Within the controlled frame-budget regime that gove