activity
20242026
most citedOne Token to Fool LLM-as-a-Judge

1 citations · 1 across the 3 of their papers we have counts for

collaborators

13 papers

cs.LG2026

Moving Out: Physically-grounded Human-AI Collaboration

Xuhui Kang, Sung-Wook Lee, Haolin Liu +2

The ability to adapt to physical actions and constraints in an environment is crucial for embodied agents (e.g., robots) to effectively collaborate with humans. Such physically gro…

cs.LG20261 cited

One Token to Fool LLM-as-a-Judge

Yulai Zhao, Haolin Liu, Dian Yu +4

Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-ba…

cs.LG2026

On the Complexity of Offline Reinforcement Learning with -Approximation and Partial Coverage

Haolin Liu, Braham Snyder, Chen-Yu Wei

We study offline reinforcement learning under -approximation and partial coverage, a setting that motivates practical algorithms such as Conservative -Learning (CQL; Ku…

cs.LG2026

An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs

Haolin Liu, Chen-Yu Wei, Julian Zimmert

We study decision making with structured observation (DMSO). Previous work (Foster et al., 2021b, 2023a) has characterized the complexity of DMSO via the decision-estimation coeffi…

cs.LG2026

Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation

Yujun Zhou, Zhenwen Liang, Haolin Liu +7

Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve wit…

cs.LG2026

Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning

Haolin Liu, Dian Yu, Sidi Lu +6

Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…