7 papers
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
Xuanhua He, Quande Liu, Zixuan Ye +7
Fine-grained and efficient controllability on video diffusion transformers has raised increasing desires for the applicability. Recently, In-context Conditioning emerged as a power…
UNIC: Unified In-Context Video Editing
Zixuan Ye, Xuanhua He, Quande Liu +7
Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional ad…
Multi-party Collaborative Attention Control for Image Customization
Han Yang, Chuanguang Yang, Qiuli Wang +4
The rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically acce…
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
Xuan Ju, Weicai Ye, Quande Liu +6
Current video generative foundation models primarily focus on text-to-video tasks, providing limited control for fine-grained video content creation. Although adapter-based approac…
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
Zhichao Liao, Xiaokun Liu, Wenyu Qin +6
Image Aesthetic Assessment (IAA) is a long-standing and challenging research task. However, its subset, Human Image Aesthetic Assessment (HIAA), has been scarcely explored. To brid…
Improving Video Generation with Human Feedback
Jie Liu, Gongye Liu, Jiajun Liang +14
Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this w…