collaborators

12 papers

cs.LG2026

Regret minimization in Linear Bandits with offline data via extended D-optimal exploration

Sushant Vijayan, Arun Suggala, Karthikeyan Shanmugam +1

We consider the problem of online regret minimization in linear bandits with access to prior observations (offline data) from the underlying bandit model. There are numerous applic…

cs.AI2025

Learning to Call: A Field Trial of a Collaborative Bandit Algorithm for Improved Message Delivery in Mobile Maternal Health

Arpan Dasgupta, Mizhaan Maniyar, Awadhesh Srivastava +7

Mobile health (mHealth) programs utilize automated voice messages to deliver health information, particularly targeting underserved communities, demonstrating the effectiveness of…

cs.RO2025

ROPES: Robotic Pose Estimation via Score-Based Causal Representation Learning

Pranamya Kulkarni, Puranjay Datta, Burak Varıcı +3

Causal representation learning (CRL) has emerged as a powerful unsupervised framework that (i) disentangles the latent generative factors underlying high-dimensional data, and (ii)…

cs.LG2025

A Personalized Exercise Assistant using Reinforcement Learning (PEARL): Results from a four-arm Randomized-controlled Trial

Amy Armento Lee, Narayan Hegde, Nina Deliu +16

Consistent physical inactivity poses a major global health challenge. Mobile health (mHealth) interventions, particularly Just-in-Time Adaptive Interventions (JITAIs), offer a prom…

cs.LG2025

Score-based Causal Representation Learning: Linear and General Transformations

Burak Varıcı, Emre Acartürk, Karthikeyan Shanmugam +2

This paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transformation that maps the latent va…

cs.LG2025

Robust Reward Modeling via Causal Rubrics

Pragya Srivastava, Harman Singh, Rahul Madhavan +9

Reward models (RMs) are fundamental to aligning Large Language Models (LLMs) via human feedback, yet they often suffer from reward hacking. They tend to latch on to superficial or…