12 papers
Regret minimization in Linear Bandits with offline data via extended D-optimal exploration
Sushant Vijayan, Arun Suggala, Karthikeyan Shanmugam +1
We consider the problem of online regret minimization in linear bandits with access to prior observations (offline data) from the underlying bandit model. There are numerous applic…
Learning to Call: A Field Trial of a Collaborative Bandit Algorithm for Improved Message Delivery in Mobile Maternal Health
Arpan Dasgupta, Mizhaan Maniyar, Awadhesh Srivastava +7
Mobile health (mHealth) programs utilize automated voice messages to deliver health information, particularly targeting underserved communities, demonstrating the effectiveness of…
ROPES: Robotic Pose Estimation via Score-Based Causal Representation Learning
Pranamya Kulkarni, Puranjay Datta, Burak Varıcı +3
Causal representation learning (CRL) has emerged as a powerful unsupervised framework that (i) disentangles the latent generative factors underlying high-dimensional data, and (ii)…
A Personalized Exercise Assistant using Reinforcement Learning (PEARL): Results from a four-arm Randomized-controlled Trial
Amy Armento Lee, Narayan Hegde, Nina Deliu +16
Consistent physical inactivity poses a major global health challenge. Mobile health (mHealth) interventions, particularly Just-in-Time Adaptive Interventions (JITAIs), offer a prom…
Score-based Causal Representation Learning: Linear and General Transformations
Burak Varıcı, Emre Acartürk, Karthikeyan Shanmugam +2
This paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transformation that maps the latent va…
Robust Reward Modeling via Causal Rubrics
Pragya Srivastava, Harman Singh, Rahul Madhavan +9
Reward models (RMs) are fundamental to aligning Large Language Models (LLMs) via human feedback, yet they often suffer from reward hacking. They tend to latch on to superficial or…