attention mechanisms 2efficient inference 1linear attention 1Mamba architecture 1spatial-temporal modeling 1token localization 1transformers 1video classification 1vision models 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.CV2026
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Fanghui Xue +5
The paper introduces VideoSEMA, a split space‑time attention model for video classification that combines a scalable Mamba‑like spatial attention block with softmax temporal attent…
cs.CV2026
SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4
The paper introduces SEMA, a new attention mechanism for vision transformers that combines token localization with arithmetic averaging to avoid the dispersion problem of linear at…