4 papers
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker
End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure…
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
Xiaoyang Cao, Zelai Xu, Mo Guang +4
Standard human preference-based alignment methods, such as Reinforcement Learning from Human Feedback (RLHF), are a cornerstone for aligning large language models (LLMs) with human…
Pareto Control Barrier Function for Inner Safe Set Maximization Under Input Constraints
Xiaoyang Cao, Zhe Fu, Alexandre M. Bayen
This article introduces the Pareto Control Barrier Function (PCBF) algorithm to maximize the inner safe set of dynamical systems under input constraints. Traditional Control Barrie…
Virtual Nodes Improve Long-term Traffic Prediction
Xiaoyang Cao, Dingyi Zhuang, Jinhua Zhao +1
Effective traffic prediction is a cornerstone of intelligent transportation systems, enabling precise forecasts of traffic flow, speed, and congestion. While traditional spatio-tem…