3 papers
cs.CV2026
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
Bingqi Ma, Linlong Lang, Ming Zhang +5
The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusio…
cs.CV2025
Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting
Xingzhong Hou, Jie Wu, Boxiao Liu +5
Image inpainting is the task of reconstructing missing or damaged parts of an image in a way that seamlessly blends with the surrounding content. With the advent of advanced genera…
cs.CV2024
See Further When Clear: Curriculum Consistency Model
Yunpeng Liu, Boxiao Liu, Yi Zhang +4
Significant advances have been made in the sampling efficiency of diffusion models and flow matching models, driven by Consistency Distillation (CD), which trains a student model t…