Skip to content
arXiv cs.CV · Papers

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

arXiv:2606.25842v2 Announce Type: replace Abstract: Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations. As a result, current video understanding approaches, especially in egocentric settings characterized by complex dyn