4 papers
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
Xiaomeng Yang, Mengping Yang, Junyan Wang +3
Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual generation. However, existing alignmen…
SARA: Structural and Adversarial Representation Alignment for Training-efficient Diffusion Models
Hesen Chen, Junyan Wang, Zhiyu Tan +1
Modern diffusion models encounter a fundamental trade-off between training efficiency and generation quality. While existing representation alignment methods, such as REPA, acceler…
Raccoon: Multi-stage Diffusion Training with Coarse-to-Fine Curating Videos
Zhiyu Tan, Junyan Wang, Hao Yang +4
Text-to-video generation has demonstrated promising progress with the advent of diffusion models, yet existing approaches are limited by dataset quality and computational resources…
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
Yibin Wang, Zhiyu Tan, Junyan Wang +3
Recent advances in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos with human pr…