1 paper
You Qin, Linqing Wang, Hao Fei +4
The post-training pipeline for diffusion models currently has two stages: supervised fine-tuning (SFT) on curated data and reinforcement learning (RL) with reward models. A fundame…