5 papers · 1 filter
Why DPO is a Misspecified Estimator and How to Fix It
Aditya Gopalan, Sayak Ray Chowdhury, Debangshu Banerjee
Direct alignment algorithms such as Direct Preference Optimization (DPO) fine-tune models based on preference data, using only supervised learning instead of two-stage reinforcemen…
Testing the Feasibility of Linear Programs with Bandit Feedback
Aditya Gangrade, Aditya Gopalan, Venkatesh Saligrama +1
While the recent literature has seen a surge in the study of constrained bandit problems, all existing methods for these begin by assuming the feasibility of the underlying problem…
A Unified Framework for Discovering Discrete Symmetries
Pavan Karjol, Rohan Kashyap, Aditya Gopalan +1
We consider the problem of learning a function respecting a symmetry from among a class of symmetries. We develop a unified framework that enables symmetry discovery across a broad…
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
Debangshu Banerjee, Aditya Gopalan
Parametric, feature-based reward models are employed by a variety of algorithms in decision-making settings such as bandits and Markov decision processes (MDPs). The typical assump…
Online Learning in Kernelized Markov Decision Processes
Sayak Ray Chowdhury, Aditya Gopalan
We consider online learning for minimizing regret in unknown, episodic Markov decision processes (MDPs) with continuous states and actions. We develop variants of the UCRL and post…