Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
Haojie Zhang, Zhihao Liang, Ruibo Fu +5
Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and computational efficiency. As video le…
cs.CV2025
A Separable Self-attention Inspired by the State Space Model for Computer Vision
Juntao Zhang, Shaogeng Liu, Kun Bian +5
Mamba is an efficient State Space Model (SSM) with linear computational complexity. Although SSMs are not suitable for handling non-causal data, Vision Mamba (ViM) methods still de…