3 papers
cs.LG2026
Displacement-Resistant Extensions of DPO with Nonconvex -Divergences
Idan Pipano, Shoham Sabach, Kavosh Asadi +1
DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference…
cs.LG2025
C2-DPO: Constrained Controlled Direct Preference Optimization
Kavosh Asadi, Julien Han, Idan Pipano +5
Direct preference optimization (\texttt{DPO}) has emerged as a promising approach for solving the alignment problem in AI. In this paper, we make two counter-intuitive observations…
cs.GT2025
On the Convergence of No-Regret Dynamics in Information Retrieval Games with Proportional Ranking Functions
Omer Madmon, Idan Pipano, Itamar Reinman +1
Publishers who publish their content on the web act strategically, in a behavior that can be modeled within the online learning framework. Regret, a central concept in machine lear…