20 citations · 20 across the 4 of their papers we have counts for
3 papers · 1 filter
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker
End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure…
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
Xiaoyang Cao, Zelai Xu, Mo Guang +4
Standard human preference-based alignment methods, such as Reinforcement Learning from Human Feedback (RLHF), are a cornerstone for aligning large language models (LLMs) with human…
Robust Model-based Reinforcement Learning for Autonomous Greenhouse Control
Wanpeng Zhang, Xiaoyan Cao, Yao Yao +3
Due to the high efficiency and less weather dependency, autonomous greenhouses provide an ideal solution to meet the increasing demand for fresh food. However, managers are faced w…