7 papers
Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-Tuning
Shresth Verma, Mauricio Tec, Cheol Woo Kim +2
While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a scalable approach, yet its perfor…
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
Shresth Verma, Arpan Dasgupta, Neha Madhiwalla +2
Maternal and child health is a critical concern around the world. In many global health programs disseminating preventive care and health information, limited healthcare worker res…
Preference Robustness for DPO with Applications to Public Health
Cheol Woo Kim, Shresth Verma, Mauricio Tec +1
We study an LLM fine-tuning task for designing reward functions for sequential resource allocation problems in public health, guided by human preferences expressed in natural langu…
Lightweight Robust Direct Preference Optimization
Cheol Woo Kim, Shresth Verma, Mauricio Tec +1
Direct Preference Optimization (DPO) has become a popular method for fine-tuning large language models (LLMs) due to its stability and simplicity. However, it is also known to be s…
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
Shresth Verma, Niclas Boehmer, Lingkai Kong +1
LLMs are increasingly used to design reward functions based on human preferences in Reinforcement Learning (RL). We focus on LLM-designed rewards for Restless Multi-Armed Bandits,…
Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning
Cheol Woo Kim, Jai Moondra, Shresth Verma +4
In many real-world applications of reinforcement learning (RL), deployed policies have varied impacts on different stakeholders, creating challenges in reaching consensus on how to…