Skip to content
arXiv cs.CV · Papers

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning

arXiv:2607.19547v1 Announce Type: new Abstract: Long-video question answering requires a model to preserve visual evidence over time without repeatedly reprocessing the same video. A practical approach is to store the vision-language model's internal key-value (KV) cache for each video chunk and retrieve that state at