collaborators

6 papers

cs.LG2026

Stabilizing Policy Optimization via Logits Convexity

Hongzhan Chen, Tao Yang, Yuhua Zhu +3

While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to super…

cs.LG2025

Merge and Guide: Unifying Model Merging and Guided Decoding for Controllable Multi-Objective Generation

Guofu Xie, Chen Zhang, Xiao Zhang +3

Adapting to diverse user needs at test time is a key challenge in controllable multi-objective generation. Existing methods are insufficient: merging-based approaches provide indir…

cs.AI2025

From <Answer> to <Think>: Multidimensional Supervision of Reasoning Process for LLM Optimization

Beining Wang, Weihang Su, Hongtao Tian +5

Improving the multi-step reasoning ability of Large Language Models (LLMs) is a critical yet challenging task. The dominant paradigm, outcome-supervised reinforcement learning (RLV…

cs.LG2025

Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization

Chao Wang, Tao Yang, Hongtao Tian +5

Critic-free methods like GRPO reduce memory demands by estimating advantages from multiple rollouts but tend to converge slowly, as critical learning signals are diluted by an abun…

cs.LG2025

Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation

Guofu Xie, Xiao Zhang, Ting Yao +1

User information needs are often highly diverse and varied. A key challenge in current research is how to achieve controllable multi-objective generation while enabling rapid adapt…

cs.CL2025

Discriminative Policy Optimization for Token-Level Reward Models

Hongzhan Chen, Tao Yang, Shiping Gao +4

Process reward models (PRMs) provide more nuanced supervision compared to outcome reward models (ORMs) for optimizing policy models, positioning them as a promising approach to enh…