1 citations · 1 across the 17 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models
Jade Zou, Tao Huang, Weijie Kong +7
Reinforcement learning (RL) has become an effective way to improve prompt alignment and perceptual quality in diffusion and flow-matching generators. A critical step for applying o…
cs.LG2026
SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models
You Qin, Linqing Wang, Hao Fei +4
The post-training pipeline for diffusion models currently has two stages: supervised fine-tuning (SFT) on curated data and reinforcement learning (RL) with reward models. A fundame…