4 papers
Policy Improvement Reinforcement Learning
Huaiyang Wang, Xiaojie Li, Xiaohan Wang +10
Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they c…
Constrained Language Model Policy Optimization via Risk-aware Stepwise Alignment
Lijun Zhang, Lin Li, Wei Wei +5
When fine-tuning pre-trained Language Models (LMs) to exhibit desired behaviors, maintaining control over risk is critical for ensuring both safety and trustworthiness. Most existi…
Risk-aware Direct Preference Optimization under Nested Risk Measure
Lijun Zhang, Lin Li, Yajie Qi +4
When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also i…
Mean Field Correlated Imitation Learning
Zhiyu Zhao, Qirui Mi, Ning Yang +4
We investigate multi-agent imitation learning (IL) within the framework of mean field games (MFGs), considering the presence of time-varying correlated signals. Existing MFG IL alg…