8 papers
REVES: REvision and VErification--Augmented Training for Test-Time Scaling
Yuanxin Liu, Ruida Zhou, Xinyan Zhao +6
Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM) reasoning. However, standard post-training methods primarily o…
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
Amirhossein Afsharrad, Ruida Zhou, Luca Viano +2
Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal prefere…
Displacement-Resistant Extensions of DPO with Nonconvex -Divergences
Idan Pipano, Shoham Sabach, Kavosh Asadi +1
DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference…
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
Batuhan K. Karaman, Aditya Rawal, Suhaila Shakiah +4
Reinforcement learning with verifiable rewards has emerged as a promising paradigm for enhancing the reasoning capabilities of large language models particularly in mathematics. Cu…
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
Luca Viano, Ruida Zhou, Yifan Sun +4
The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…
Directional-Clamp PPO
Gilad Karpel, Ruida Zhou, Shoham Sabach +1
Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a rang…