7 papers
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics
Vedant Palit, Udvas Das, Brahim Driss +1
As AI-driven Decision Makers (ADMs) influence our socioeconomic reality, their roles in both enhancing efficiency and amplifying the social biases have drawn attention. In this pap…
Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts
Udvas Das, Waris Radji, Debabrota Basu +1
We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized pref…
Learning to Explore with Lagrangians for Bandits under Unknown Linear Constraints
Udvas Das, Debabrota Basu
Pure exploration in bandits formalises multiple real-world problems, such as tuning hyper-parameters or conducting user studies to test a set of items, where different safety, reso…
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
Debabrota Basu, Udvas Das, Brahim Driss +1
Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) metho…
Witness Set in Monotone Polygons: Exact and Approximate
Udvas Das, Binayak Dutta, Satyabrata Jana +2
Given a simple polygon , two points and within are {\em visible} to each other if the line segment between and is contained in $\mathscr{…
FraPPE: Fast and Efficient Preference-based Pure Exploration
Udvas Das, Apurv Shukla, Debabrota Basu
Preference-based Pure Exploration (PrePEx) aims to identify with a given confidence level the set of Pareto optimal arms in a vector-valued (aka multi-objective) bandit, where the…