collaborators

5 papers

cs.LG2026

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Wenlong Deng, Jiaji Huang, Kaan Ozkara +4

Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinfor…

cs.AI2026

AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

Yuyang Hu, Hongjin Qian, Shuting Wang +5

Recent progress on long-horizon agentic tasks has been driven largely by scaling up individual agents through stronger models, better tools, and more effective scaffolding. In cont…

cs.LG2026

Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning

Wenlong Deng, Yi Ren, Yushu Li +4

Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward explorati…

cs.CL2026

On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement

Wenlong Deng, Yushu Li, Boying Gong +3

Tool-integrated (TI) reinforcement learning (RL) enables large language models (LLMs) to perform multi-step reasoning by interacting with external tools such as search engines and…

cs.LG2025

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Wenlong Deng, Yi Ren, Muchen Li +3

Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a…