collaborators

6 papers

cs.CL2026

TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles

Yirong Zeng, Yufei Liu, Xiao Ding +9

Instruction Following (IF) is a core capability of LLMs, requiring strict adherence to diverse constraints, ranging from verifiable ones (e.g., output length) to unverifiable ones…

cs.LG2026

TreeAdv: Tree-Structured Advantage Redistribution for Group-Based RL

Lang Cao, Hui Ruan, Yongqian Li +5

Reinforcement learning with group-based objectives, such as Group Relative Policy Optimization (GRPO), is a common framework for aligning large language models on complex reasoning…

cs.AI2026

AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints

Yirong Zeng, Xiao Ding, Yufei Liu +9

Tool use represents a critical capability for AI agents, with recent advances focusing on leveraging reinforcement learning (RL) to scale up the explicit reasoning process to achie…

cs.AI2026

The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?

Yirong Zeng, Shen You, Yufei Liu +9

Equipping LLMs with external tools effectively addresses internal reasoning limitations. However, it introduces a critical yet under-explored phenomenon: tool overuse, the unnecess…

cs.LG2026

IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models

Haonan Song, Qingchen Xie, Huan Zhu +12

Generative Reward Models (GRMs) have demonstrated strong performance in reward modeling, due to their interpretability and potential for refinement through reinforcement learning (…

cs.LG2026

Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following

Yirong Zeng, Yufei Liu, Xiao Ding +9

A central belief in scaling reinforcement learning with verifiable rewards for instruction following (IF) tasks is that, a diverse mixture of verifiable hard and unverifiable soft…