HF Daily Papers
· Papers
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, quer