3 papers
cs.LG2026
Learning Markov Decision Processes under Fully Bandit Feedback
Zhengjia Zhuo, Anupam Gupta, Viswanath Nagarajan
A standard assumption in Reinforcement Learning is that the agent observes every visited state-action pair in the associated Markov Decision Process (MDP), along with the per-step…
cs.DS2025
The Average-Value Allocation Problem
Kshipra Bhawalkar, Zhe Feng, Anupam Gupta +3
We initiate the study of centralized algorithms for welfare-maximizing allocation of goods to buyers subject to average-value constraints. We show that this problem is NP-hard to a…
cs.DS2025
A Little Clairvoyance Is All You Need
Anupam Gupta, Haim Kaplan, Alexander Lindermayr +2
We revisit the classical problem of minimizing the total flow time of jobs on a single machine in the online setting where jobs arrive over time. It has long been known that the Sh…