Skip to content
Apple ML Research · Cloud & Big Tech

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark for evaluating long-for