Skip to content
arXiv cs.CV · Papers

A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models

arXiv:2503.15846v2 Announce Type: replace Abstract: Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graphs remain limited compared to the rapid advances in Multimodal Large Language Models (MLL