3 papers
stat.ML2025
Reinforcement Learning in MDPs with Information-Ordered Policies
Zhongjun Zhang, Shipra Agrawal, Ilan Lobel +2
We propose an epoch-based reinforcement learning algorithm for infinite-horizon average-cost Markov decision processes (MDPs) that leverages a partial order over a policy class. In…
stat.ML2025
Tensor Completion with Nearly Linear Samples Given Weak Side Information
Christina Lee Yu, Xumei Xi
Tensor completion exhibits an interesting computational-statistical gap in terms of the number of samples needed to perform tensor estimation. While there are only degrees…
cs.LG2025
Artificial Replay: A Meta-Algorithm for Harnessing Historical Data in Bandits
Siddhartha Banerjee, Sean R. Sinclair, Milind Tambe +2
Most real-world deployments of bandit algorithms exist somewhere in between the offline and online set-up, where some historical data is available upfront and additional data is co…