5 papers
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
Shaobo Wang, Yujie Chen, Yafeng Sun +9
Modern reasoning agents are increasingly evaluated on their ability to generate multiple valid solution paths, plans, or tool-use traces for a given input. Standard reward-maximizi…
Explainable reinforcement learning from human feedback to improve alignment
Shicheng Liu, Siyuan Xu, Wenjie Qiu +2
A common and effective strategy for humans to improve an unsatisfactory outcome in daily life is to find a cause of this outcome and correct the cause. In this paper, we investigat…
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
Wenjie Qiu, Yi-Chen Li, Xuqin Zhang +4
Learning reward models from human preference datasets and subsequently optimizing language models via reinforcement learning has emerged as a fundamental paradigm for aligning LLMs…
Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics
Xinyu Zhang, Wenjie Qiu, Yi-Chen Li +4
Developing policies that can adjust to non-stationary environments is essential for real-world reinforcement learning applications. However, learning such adaptable policies in off…
Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation
Yi-Chen Li, Fuxiang Zhang, Wenjie Qiu +5
Large Language Models (LLMs), trained on a large amount of corpus, have demonstrated remarkable abilities. However, it may not be sufficient to directly apply open-source LLMs like…