collaborators

5 papers

cs.LG2026

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL

Shaobo Wang, Yujie Chen, Yafeng Sun +9

Modern reasoning agents are increasingly evaluated on their ability to generate multiple valid solution paths, plans, or tool-use traces for a given input. Standard reward-maximizi…

cs.LG2025

Explainable reinforcement learning from human feedback to improve alignment

Shicheng Liu, Siyuan Xu, Wenjie Qiu +2

A common and effective strategy for humans to improve an unsatisfactory outcome in daily life is to find a cause of this outcome and correct the cause. In this paper, we investigat…

cs.CL2025

Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference

Wenjie Qiu, Yi-Chen Li, Xuqin Zhang +4

Learning reward models from human preference datasets and subsequently optimizing language models via reinforcement learning has emerged as a fundamental paradigm for aligning LLMs…

cs.LG2025

Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics

Xinyu Zhang, Wenjie Qiu, Yi-Chen Li +4

Developing policies that can adjust to non-stationary environments is essential for real-world reinforcement learning applications. However, learning such adaptable policies in off…

cs.LG2025

Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation

Yi-Chen Li, Fuxiang Zhang, Wenjie Qiu +5

Large Language Models (LLMs), trained on a large amount of corpus, have demonstrated remarkable abilities. However, it may not be sufficient to directly apply open-source LLMs like…