2 papers
cs.LG2026
Learning Markov Decision Processes under Fully Bandit Feedback
Zhengjia Zhuo, Anupam Gupta, Viswanath Nagarajan
A standard assumption in Reinforcement Learning is that the agent observes every visited state-action pair in the associated Markov Decision Process (MDP), along with the per-step…
cs.DS2025
A Little Clairvoyance Is All You Need
Anupam Gupta, Haim Kaplan, Alexander Lindermayr +2
We revisit the classical problem of minimizing the total flow time of jobs on a single machine in the online setting where jobs arrive over time. It has long been known that the Sh…