4 papers
Simple and Robust Quality Disclosure: The Power of Quantile Partition
Shipra Agrawal, Yiding Feng, Wei Tang
Quality information on online platforms is often conveyed through simple, percentile-based badges and tiers that remain stable across different market environments. Motivated by th…
Q-learning with Posterior Sampling
Priyank Agrawal, Shipra Agrawal, Azmat Azati
Bayesian posterior sampling techniques have demonstrated superior empirical performance in many exploration-exploitation settings. However, their theoretical analysis remains a cha…
Reinforcement Learning in MDPs with Information-Ordered Policies
Zhongjun Zhang, Shipra Agrawal, Ilan Lobel +2
We propose an epoch-based reinforcement learning algorithm for infinite-horizon average-cost Markov decision processes (MDPs) that leverages a partial order over a policy class. In…
Optimistic Q-learning for average reward and episodic reinforcement learning
Priyank Agrawal, Shipra Agrawal
We present an optimistic Q-learning algorithm for regret minimization in average reward reinforcement learning under an additional assumption on the underlying MDP that for all pol…