arXiv cs.CV
· Papers
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
arXiv:2607.24794v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding. To accommodate this constraint, models typically resort to keyframe selection. However, unifor