VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
arXiv:2607.14711v2 Announce Type: replace Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA)…