4 papers
Probe-then-Commit Multi-Objective Bandits: Theoretical Benefits of Limited Multi-Arm Feedback
Ming Shi
We study an online resource-selection problem motivated by multi-radio access selection and mobile edge computing offloading. In each round, an agent chooses among candidate li…
Bi-Level Online Provisioning and Scheduling with Switching Costs and Cross-Level Constraints
Jialei Liu, C. Emre Koksal, Ming Shi
We study a bi-level online provisioning and scheduling problem motivated by network resource allocation, where provisioning decisions are made at a slow time scale while queue-/sta…
Communication-Corruption Coupling and Verification in Cooperative Multi-Objective Bandits
Ming Shi
We study cooperative stochastic multi-armed bandits with vector-valued rewards under adversarial corruption and limited verification. In each of rounds, each of agents sele…
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
Ming Shi, Yingbin Liang, Ness B. Shroff
Partially observable Markov decision processes (POMDPs) are a general framework for sequential decision-making under latent state uncertainty, yet learning in POMDPs is intractable…