Skip to content
arXiv cs.CV · Papers

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

arXiv:2608.04302v1 Announce Type: new Abstract: Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video descr