collaborators

8 papers

cs.LG2026

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

Yuanxin Liu, Ruida Zhou, Xinyan Zhao +6

Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM) reasoning. However, standard post-training methods primarily o…

cs.LG2026

Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback

Amirhossein Afsharrad, Ruida Zhou, Luca Viano +2

Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal prefere…

cs.LG2026

Displacement-Resistant Extensions of DPO with Nonconvex -Divergences

Idan Pipano, Shoham Sabach, Kavosh Asadi +1

DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference…

cs.CL2026

DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning

Batuhan K. Karaman, Aditya Rawal, Suhaila Shakiah +4

Reinforcement learning with verifiable rewards has emerged as a promising paradigm for enhancing the reasoning capabilities of large language models particularly in mathematics. Cu…

cs.LG2026

Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains

Luca Viano, Ruida Zhou, Yifan Sun +4

The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…

cs.LG2025

Directional-Clamp PPO

Gilad Karpel, Ruida Zhou, Shoham Sabach +1

Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a rang…