8 papers
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
Anvay Shah, Ramsundar Anandanarayanan, Sharayu Moharir +1
A Tree Markov Decision Problem (T-MDP) is a finite-horizon MDP with a starting state , in which every state is reachable from through exactly one state-action trajec…
Cascading Bandits With Feedback
R Sri Prakash, Nikhil Karamchandani, Sharayu Moharir
Motivated by the challenges of edge inference, we study a variant of the cascade bandit model in which each arm corresponds to an inference model with an associated accuracy and er…
Fixed-Budget Constrained Best Arm Identification in Grouped Bandits
Raunak Mukherjee, Sharayu Moharir
We study fixed budget constrained best-arm identification in grouped bandits, where each arm consists of multiple independent attributes with stochastic rewards. An arm is consider…
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
Vishnu Narayanan Moothedath, Umang Agarwal, Umeshraja N +3
We focus on a binary classification problem in an edge intelligence system where false negatives are more costly than false positives. The system has a compact, locally deployed mo…
Low-Regret and Low-Complexity Learning for Hierarchical Inference
Sameep Chattopadhyay, Vinay Sutar, Jaya Prakash Champati +1
This work focuses on Hierarchical Inference (HI) in edge intelligence systems, where a compact Local-ML model on an end-device works in conjunction with a high-accuracy Remote-ML m…
Observation-Free Attacks on Online Learning to Rank
Sameep Chattopadhyay, Nikhil Karamchandani, Sharayu Moharir
Online learning to rank (OLTR) plays a critical role in information retrieval and machine learning systems, with a wide range of applications in search engines and content recommen…