3 papers
cs.LG2026
Learning Markov Decision Processes under Fully Bandit Feedback
Zhengjia Zhuo, Anupam Gupta, Viswanath Nagarajan
A standard assumption in Reinforcement Learning is that the agent observes every visited state-action pair in the associated Markov Decision Process (MDP), along with the per-step…
cs.DS2025
A Simple Approximation Algorithm for Optimal Decision Tree
Zhengjia Zhuo, Viswanath Nagarajan
Optimal decision tree (\odt) is a fundamental problem arising in applications such as active learning, entity identification, and medical diagnosis. An instance of \odt is given by…
cs.DS2025
Identifying Approximate Minimizers under Stochastic Uncertainty
Hessa Al-Thani, Viswanath Nagarajan
We study a fundamental stochastic selection problem involving independent random variables, each of which can be queried at some cost. Given a tolerance level , the goal is…