3 papers
cs.CV2024
QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model
Fei Xie, Weijia Zhang, Zhongdao Wang +1
Recent advancements in State Space Models, notably Mamba, have demonstrated superior performance over the dominant Transformer models, particularly in reducing the computational co…
cs.CV2024
Video Prediction Transformers without Recurrence or Convolution
Yujin Tang, Lu Qi, Xiangtai Li +2
Video prediction has witnessed the emergence of RNN-based models led by ConvLSTM, and CNN-based models led by SimVP. Following the significant success of ViT, recent works have int…
cs.CV2024
Correlation-Embedded Transformer Tracking: A Single-Branch Framework
Fei Xie, Wankou Yang, Chunyu Wang +4
Developing robust and discriminative appearance models has been a long-standing research challenge in visual object tracking. In the prevalent Siamese-based paradigm, the features…