arXiv cs.AI
· Papers
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
arXiv:2603.18558v2 Announce Type: replace-cross Abstract: Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows. Within the controlled frame-budget regime that gove