activity
20242026
collaborators

9 papers

cs.LG2026

Displacement-Resistant Extensions of DPO with Nonconvex -Divergences

Idan Pipano, Shoham Sabach, Kavosh Asadi +1

DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference…

cs.LG2026

Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains

Luca Viano, Ruida Zhou, Yifan Sun +4

The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…

cs.LG2025

Directional-Clamp PPO

Gilad Karpel, Ruida Zhou, Shoham Sabach +1

Proximal Policy Optimization (PPO) is widely regarded as one of the most successful deep reinforcement learning algorithms, known for its robustness and effectiveness across a rang…

cs.LG2025

ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs

Hongyi Liu, Rajarshi Saha, Zhen Jia +5

Large Language Models (LLMs) have demonstrated exceptional performance in natural language processing tasks, yet their massive size makes serving them inefficient and costly. Semi-…

cs.LG2025

C2-DPO: Constrained Controlled Direct Preference Optimization

Kavosh Asadi, Julien Han, Idan Pipano +5

Direct preference optimization (\texttt{DPO}) has emerged as a promising approach for solving the alignment problem in AI. In this paper, we make two counter-intuitive observations…

math.OC2025

Dynamic FISTA for Convex Composite Bi-Level Optimization

Roey Merchav, Shoham Sabach, Marc Teboulle

In this paper, we study convex bi-level optimization problems where both the inner and outer levels are given as a composite convex minimization. We propose the Fast Bi-level Proxi…