11 papers
Moving Out: Physically-grounded Human-AI Collaboration
Xuhui Kang, Sung-Wook Lee, Haolin Liu +2
The ability to adapt to physical actions and constraints in an environment is crucial for embodied agents (e.g., robots) to effectively collaborate with humans. Such physically gro…
One Token to Fool LLM-as-a-Judge
Yulai Zhao, Haolin Liu, Dian Yu +4
Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-ba…
On the Complexity of Offline Reinforcement Learning with -Approximation and Partial Coverage
Haolin Liu, Braham Snyder, Chen-Yu Wei
We study offline reinforcement learning under -approximation and partial coverage, a setting that motivates practical algorithms such as Conservative -Learning (CQL; Ku…
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We study decision making with structured observation (DMSO). Previous work (Foster et al., 2021b, 2023a) has characterized the complexity of DMSO via the decision-estimation coeffi…
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
Yujun Zhou, Zhenwen Liang, Haolin Liu +7
Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve wit…
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
Haolin Liu, Dian Yu, Sidi Lu +6
Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…