collaborators

6 papers

cs.LG2026

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

Yuanxin Liu, Ruida Zhou, Xinyan Zhao +6

Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM) reasoning. However, standard post-training methods primarily o…

cs.LG2026

HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents

Jiangweizhi Peng, Yuanxin Liu, Ruida Zhou +4

Training LLMs as interactive agents for multi-turn decision-making remains challenging, particularly in long-horizon tasks with sparse and delayed rewards, where agents must execut…

cs.LG2026

Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback

Amirhossein Afsharrad, Ruida Zhou, Luca Viano +2

Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal prefere…

cs.CL2026

DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning

Batuhan K. Karaman, Aditya Rawal, Suhaila Shakiah +4

Reinforcement learning with verifiable rewards has emerged as a promising paradigm for enhancing the reasoning capabilities of large language models particularly in mathematics. Cu…

cs.LG2026

Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains

Luca Viano, Ruida Zhou, Yifan Sun +4

The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…

cs.LG2025

Directional-Clamp PPO

Gilad Karpel, Ruida Zhou, Shoham Sabach +1

Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a rang…