24 citations · 27 across the 4 of their papers we have counts for
4 papers
Low-rank Bandits with Latent Mixtures
Aditya Gopalan, Odalric-Ambrym Maillard, Mohammadi Zaki
We study the task of maximizing rewards from recommending items (actions) to users sequentially interacting with a recommender system. Users are modeled as latent mixtures of C man…
Optimal WiFi Sensing via Dynamic Programming
Abhinav Kumar, Rahul Vaze, Sibi Raj B Pillai +1
The problem of finding an optimal sensing schedule for a mobile device that encounters an intermittent WiFi access opportunity is considered. At any given time, the WiFi is in any…
Thompson Sampling for Learning Parameterized Markov Decision Processes
Aditya Gopalan, Shie Mannor
We consider reinforcement learning in parameterized Markov Decision Processes (MDPs), where the parameterization may induce correlation across transition probabilities or rewards.…
Wireless Scheduling with Partial Channel State Information: Large Deviations and Optimality
Aditya Gopalan, Constantine Caramanis, Sanjay Shakkottai
We consider a server serving a time-slotted queued system of multiple packet-based flows, with exogenous packet arrivals and time-varying service rates. At each time, the server ca…