9 papers
Learning from Local Walks on Dynamic Graphs with Bandit Feedback
Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni +1
We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to…
Flickering Multi-Armed Bandits
Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni +1
We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, where accessibility of the next action is…
A Unified Framework for Locality in Scalable MARL
Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni +1
Scalable methods for networked multi-agent reinforcement learning let each agent plan using only a small neighborhood of the agent graph. This works only when the system is value-l…
Multi-Agent Lipschitz Bandits
Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni +1
We study the decentralized multi-player stochastic bandit problem over a continuous, Lipschitz-structured action space where hard collisions yield zero reward. Our objective is to…
Data Attribution in Adaptive Learning
Amit Kiran Rege
Machine learning models increasingly generate their own training data -- online bandits, reinforcement learning, and post-training pipelines for language models are leading example…
The Role of Generator Access in Autoregressive Post-Training
Amit Kiran Rege
We study how generator access constrains autoregressive post-training. The central question is whether the learner is confined to fresh root-start rollouts or can return to previou…