4 papers
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
Run Ling, Ke Cao, Jian Lu +15
Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity.…
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma, Jiasong Feng, Ke Cao +4
Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue…
U-StyDiT: Ultra-high Quality Artistic Style Transfer Using Diffusion Transformers
Zhanjie Zhang, Ao Ma, Ke Cao +6
Ultra-high quality artistic style transfer refers to repainting an ultra-high quality content image using the style information learned from the style image. Existing artistic styl…
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
Jing Wang, Ao Ma, Ke Cao +9
Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle…