arXiv cs.AI
· Papers
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive