5 papers
LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing
Jianzong Wu, Hao Lian, Jiongfan Yang +12
Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified framewo…
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
Jingyuan Zhu, Biaolong Chen, Le Zhang +3
Efficiently aligning large-scale video diffusion models with human intent requires a scalable and trajectory-aware pathway that bridges the inherent discrepancy between training no…
Towards Customized Multimodal Role-Play
Chao Tang, Jianzong Wu, Qingyu Shi +5
Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while…
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
Xianghao Kong, Qiaosong Qi, Yuanbin Wang +3
Fashion video generation aims to synthesize temporally consistent videos from reference images of a designated character. Despite significant progress, existing diffusion-based met…
FuseAnyPart: Diffusion-Driven Facial Parts Swapping via Multiple Reference Images
Zheng Yu, Yaohua Wang, Siying Cui +3
Facial parts swapping aims to selectively transfer regions of interest from the source image onto the target image while maintaining the rest of the target image unchanged. Most st…