collaborators

9 papers

cs.LG2026

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Beining Wang, Weihang Su, Hongtao Tian +8

Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become…

cs.AI2026

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

Chao Wang, Hongtao Tian, Tao Yang +3

Group Relative Policy Optimization (GRPO) is a default recipe for process-supervised reinforcement learning of LLM reasoners, and dense process supervision -- via learned process r…

cs.LG2025

CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment

Guofu Xie, Yunsheng Shi, Hongtao Tian +2

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of Large Language Models (LLMs) by using rule-based binary feedback. However, current RLV…

cs.AI2025

From <Answer> to <Think>: Multidimensional Supervision of Reasoning Process for LLM Optimization

Beining Wang, Weihang Su, Hongtao Tian +5

Improving the multi-step reasoning ability of Large Language Models (LLMs) is a critical yet challenging task. The dominant paradigm, outcome-supervised reinforcement learning (RLV…

cs.CL2025

From Faithfulness to Correctness: Generative Reward Models that Think Critically

Qiyao Ma, Yunsheng Shi, Hongtao Tian +3

Through reinforcement learning with verifiable rewards (RLVR), large language models have achieved substantial progress in domains with easily verifiable outcomes, such as mathemat…

cs.LG2025

Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization

Chao Wang, Tao Yang, Hongtao Tian +5

Critic-free methods like GRPO reduce memory demands by estimating advantages from multiple rollouts but tend to converge slowly, as critical learning signals are diluted by an abun…