5 papers
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
Xingtong Ge, Yi Zhang, Yushi Huang +6
Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency d…
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
Dailan He, Guanlin Feng, Xingtong Ge +5
Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult to align via reinforcement lear…
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
Bingqi Ma, Linlong Lang, Ming Zhang +5
The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusio…
Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting
Xingzhong Hou, Jie Wu, Boxiao Liu +5
Image inpainting is the task of reconstructing missing or damaged parts of an image in a way that seamlessly blends with the surrounding content. With the advent of advanced genera…
See Further When Clear: Curriculum Consistency Model
Yunpeng Liu, Boxiao Liu, Yi Zhang +4
Significant advances have been made in the sampling efficiency of diffusion models and flow matching models, driven by Consistency Distillation (CD), which trains a student model t…