arXiv cs.CV
· Papers
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
arXiv:2607.14711v2 Announce Type: replace Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local