15 papers
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
Haopeng Li, Yitong Li, Junsong Chen +8
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse atten…
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
Junsong Chen, Jincheng Yu, Yitong Li +11
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p…
Variational Adapter for Cross-modal Similarity Representation
WenZhang Wei, Zhipeng Gui, Dehua Peng +2
The core of vision-language models lies in measuring cross-modal similarity within a unified representation space. However, most image-text matching or multi-class image classifica…
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Yuyang Zhao, Yicheng Pan, Qiyuan He +6
Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the str…
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
Wenxue Li, Jingjing Ren, Peng Zhang +4
High-resolution video generation faces a coupled bottleneck of optimization instability and prohibitive computational costs. The massive expansion of the token sequence not only bi…
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
Haoyi Zhu, Haozhe Liu, Yuyang Zhao +6
We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with pre…