1 paper
Alexis Jacq, Guillaume Couairon, Valentin De Bortoli +3
Distillation and Reinforcement Learning (RL) fine-tuning are the primary pillars of diffusion post-training. While traditionally studied in isolation, the interaction between these…