3 citations · 5 across the 11 of their papers we have counts for
9 papers · 1 filter
Data Deletion Can Help in Adaptive RL
Param Budhraja, Aditya Gangrade, Alex Olshevsky +1
Deploying reinforcement learning policies in the real world requires adapting to time-varying environments. We study this problem in the contextual Markov Decision Process (cMDP) f…
Linear Transformers Implicitly Discover Unified Numerical Algorithms
Patrick Lutz, Aditya Gangrade, Hadi Daneshmand +1
We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar predictio…
Constrained Linear Thompson Sampling
Aditya Gangrade, Venkatesh Saligrama
We study safe linear bandits (SLBs), where an agent selects actions from a convex set to maximize an unknown linear objective subject to unknown linear constraints in each round. E…
Testing the Feasibility of Linear Programs with Bandit Feedback
Aditya Gangrade, Aditya Gopalan, Venkatesh Saligrama +1
While the recent literature has seen a surge in the study of constrained bandit problems, all existing methods for these begin by assuming the feasibility of the underlying problem…
Safe Linear Bandits over Unknown Polytopes
Aditya Gangrade, Tianrui Chen, Venkatesh Saligrama
The safe linear bandit problem (SLB) is an online approach to linear programming with unknown objective and unknown roundwise constraints, under stochastic bandit feedback of rewar…
Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and Risk
Tianrui Chen, Aditya Gangrade, Venkatesh Saligrama
We investigate a natural but surprisingly unstudied approach to the multi-armed bandit problem under safety risk constraints. Each arm is associated with an unknown law on safety r…