2 papers
cs.CV2026
Video Prediction Transformers without Recurrence or Convolution
Yujin Tang, Lu Qi, Xiangtai Li +2
Video prediction has witnessed the emergence of RNN-based models led by ConvLSTM, and CNN-based models led by SimVP. Following the significant success of ViT, recent works have int…
cs.CV2024
QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model
Fei Xie, Weijia Zhang, Zhongdao Wang +1
Recent advancements in State Space Models, notably Mamba, have demonstrated superior performance over the dominant Transformer models, particularly in reducing the computational co…