9 papers
Displacement-Resistant Extensions of DPO with Nonconvex -Divergences
Idan Pipano, Shoham Sabach, Kavosh Asadi +1
DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference…
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
Luca Viano, Ruida Zhou, Yifan Sun +4
The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…
Directional-Clamp PPO
Gilad Karpel, Ruida Zhou, Shoham Sabach +1
Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a rang…
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
Hongyi Liu, Rajarshi Saha, Zhen Jia +5
Large Language Models (LLMs) have demonstrated exceptional performance in natural language processing tasks, yet their massive size makes serving them inefficient and costly. Semi-…
C2-DPO: Constrained Controlled Direct Preference Optimization
Kavosh Asadi, Julien Han, Idan Pipano +5
Direct preference optimization (\texttt{DPO}) has emerged as a promising approach for solving the alignment problem in AI. In this paper, we make two counter-intuitive observations…
Dynamic FISTA for Convex Composite Bi-Level Optimization
Roey Merchav, Shoham Sabach, Marc Teboulle
In this paper, we study convex bi-level optimization problems where both the inner and outer levels are given as a composite convex minimization. We propose the Fast Bi-level Proxi…