2 citations · 2 across the 8 of their papers we have counts for
Showing 2025Show all
3 papers · 1 filter
cs.CL2025
Inference-time Alignment in Continuous Space
Yige Yuan, Teng Xiao, Li Yunfan +5
Aligning large language models with human feedback at inference time has received increasing attention due to its flexibility. Existing methods rely on generating multiple response…
cs.CL2025
Incentivizing Strong Reasoning from Weak Supervision
Yige Yuan, Teng Xiao, Shuchang Tao +4
Large language models (LLMs) have demonstrated impressive performance on reasoning-intensive tasks, but enhancing their reasoning abilities typically relies on either reinforcement…
cs.LG2025
On a Connection Between Imitation Learning and RLHF
Teng Xiao, Yige Yuan, Mingxiao Li +2
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcem…