2 papers
cs.CV2025
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
Yunze Liu, Li Yi
Hybrid Mamba-Transformer networks have recently garnered broad attention. These networks can leverage the scalability of Transformers while capitalizing on Mamba's strengths in lon…
cs.CV2025
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
Yunze Liu, Peiran Wu, Cheng Liang +3
Recent Mamba-based architectures for video understanding demonstrate promising computational efficiency and competitive performance, yet struggle with overfitting issues that hinde…