1 citations · 1 across the 11 of their papers we have counts for
1 paper · 1 filter
Shengjun Zhang, Zhang Zhang, Chensheng Dai +1
Recent reinforcement learning has enhanced the flow matching models on human preference alignment. While stochastic sampling enables the exploration of denoising directions, existi…