12 citations · 17 across the 8 of their papers we have counts for
7 papers · 2 filters
SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to Single-GPU Acceleration of Visual Generation
Yuxi Liu, Haoyu Li, Zekun Zhang +12
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity…
CrossDistill: Balancing Quality and Diversity via Trajectory-Level Hybrid Few-Step Distillation
Yuxi Liu, Haoyu Li, Yixiang Cai +8
Few-step distillation accelerates diffusion models but must balance diversity and fidelity: trajectory-based distillation preserves mode coverage, while distribution matching sharp…
RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers
Zekun Zhang, Yixiang Cai, Yuxi Liu +9
Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes t…
Video = World + Event Stream
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persis…
Wan-Streamer v0.2: Higher Resolution, Same Latency
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Lianghua Huang, Zhi-Fan Wu, Wei Wang +22
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. W…