1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Hybrid Alignment Training for Large Language Models
Chenglong Wang, Hang Zhou, Kaiyan Chang +5
Alignment training is crucial for enabling large language models (LLMs) to cater to human intentions and preferences. It is typically performed based on two stages with different o…
cs.CL2023
ESRL: Efficient Sampling-based Reinforcement Learning for Sequence Generation
Chenglong Wang, Hang Zhou, Yimin Hu +5
Applying Reinforcement Learning (RL) to sequence generation models enables the direct optimization of long-term rewards (\textit{e.g.,} BLEU and human feedback), but typically requ…