7 papers
When Can Safe Controllers Adapt? Information before Commitment
Venkatesh Saligrama
Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself. The controller may use any causal, history-dependent rule and act differently…
Data Deletion Can Help in Adaptive RL
Param Budhraja, Aditya Gangrade, Alex Olshevsky +1
Deploying reinforcement learning policies in the real world requires adapting to time-varying environments. We study this problem in the contextual Markov Decision Process (cMDP) f…
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
Patrick Lutz, Themistoklis Haris, Arjun Chandra +2
Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study multi-class linear classification in the h…
Linear Transformers Implicitly Discover Unified Numerical Algorithms
Patrick Lutz, Aditya Gangrade, Hadi Daneshmand +1
We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar predictio…
Constrained Linear Thompson Sampling
Aditya Gangrade, Venkatesh Saligrama
We study safe linear bandits (SLBs), where an agent selects actions from a convex set to maximize an unknown linear objective subject to unknown linear constraints in each round. E…
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
Ruizhao Zhu, Venkatesh Saligrama
We propose Deep Companion Learning (DCL), a novel training method for Deep Neural Networks (DNNs) that enhances generalization by penalizing inconsistent model predictions compared…