Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
arXiv:2607.24794v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.…