3 citations · 3 across the 8 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
Jingyang Ou, Jiaqi Han, Minkai Xu +5
Reinforcement Learning (RL) has proven highly effective for autoregressive language models, but adapting these methods to diffusion large language models (dLLMs) presents fundament…
cs.CL2025
-PO: Generalizing Preference Optimization with -divergence Minimization
Jiaqi Han, Mingjian Jiang, Yuxuan Song +2
Preference optimization has made significant progress recently, with numerous methods developed to align language models with human preferences. This paper introduces -divergenc…