collaborators

10 papers

cs.CL2026

REAR: Test-time Preference Realignment through Reward Decomposition

Fuxiang Zhang, Pengcheng Wang, Chenran Li +6

Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to specific needs, they often r…

cs.LG2026

Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning

Tian Xu, Zexuan Chen, Zhilong Zhang +4

Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical w…

cs.LG2026

Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow Mechanism

Tian Xu, Chenyang Wang, Xiaochen Zhai +3

Adversarial imitation learning (AIL) achieves high-quality imitation by mitigating compounding errors inherent to behavioral cloning (BC), yet its adversarial optimization frequent…

cs.LG2026

Off-Policy Value-Based Reinforcement Learning for Large Language Models

Peng-Yuan Wang, Ziniu Li, Tian Xu +8

Improving data utilization efficiency is critical for scaling reinforcement learning (RL) for long-horizon tasks where generating trajectories is expensive. However, the dominant R…

cs.MA2025

Multi-agent In-context Coordination via Decentralized Memory Retrieval

Tao Jiang, Zichuan Lin, Lihe Li +6

Large transformer models, trained on diverse datasets, have demonstrated impressive few-shot performance on previously unseen tasks without requiring parameter updates. This capabi…

cs.CL2025

Generalist Reward Models: Found Inside Large Language Models

Yi-Chen Li, Tian Xu, Yang Yu +6

The alignment of Large Language Models (LLMs) is critically dependent on reward models trained on costly human preference data. While recent work explores bypassing this cost with…