5 papers · 1 filter
Regret minimization in Linear Bandits with offline data via extended D-optimal exploration
Sushant Vijayan, Arun Suggala, Karthikeyan Shanmugam +1
We consider the problem of online regret minimization in linear bandits with access to prior observations (offline data) from the underlying bandit model. There are numerous applic…
A Personalized Exercise Assistant using Reinforcement Learning (PEARL): Results from a four-arm Randomized-controlled Trial
Amy Armento Lee, Narayan Hegde, Nina Deliu +16
Consistent physical inactivity poses a major global health challenge. Mobile health (mHealth) interventions, particularly Just-in-Time Adaptive Interventions (JITAIs), offer a prom…
Robust Reward Modeling via Causal Rubrics
Pragya Srivastava, Harman Singh, Rahul Madhavan +9
Reward models (RMs) are fundamental to aligning Large Language Models (LLMs) via human feedback, yet they often suffer from reward hacking. They tend to latch on to superficial or…
Online Bidding under RoS Constraints without Knowing the Value
Sushant Vijayan, Zhe Feng, Swati Padmanabhan +3
We consider the problem of bidding in online advertising, where an advertiser aims to maximize value while adhering to budget and Return-on-Spend (RoS) constraints. Unlike prior wo…
Bayesian Collaborative Bandits with Thompson Sampling for Improved Outreach in Maternal Health Program
Arpan Dasgupta, Gagan Jain, Arun Suggala +3
Mobile health (mHealth) programs face a critical challenge in optimizing the timing of automated health information calls to beneficiaries. This challenge has been formulated as a…