5 papers · 1 filter
Adaptive Calibration in Non-Stationary Environments
Junyan Liu, Haipeng Luo, Lillian J. Ratliff
Making calibrated online predictions is a central challenge in modern AI systems. Much of the existing literature focuses on fully adversarial environments where outcomes may be ar…
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
Junyan Liu, Haipeng Luo, Zihan Zhang +1
We study online learning in two-player uninformed Markov games, where the opponent's actions and policies are unobserved. In this setting, Tian et al. (2021) show that achieving no…
Improved Regret and Contextual Linear Extension for Pandora's Box and Prophet Inequality
Junyan Liu, Ziyun Chen, Kun Wang +2
We study the Pandora's Box problem in an online learning setting with semi-bandit feedback. In each round, the learner sequentially pays to open up to boxes with unknown reward…
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
Junyan Liu, Lillian J. Ratliff
We study the repeated principal-agent bandit game, where the principal indirectly interacts with the unknown environment by proposing incentives for the agent to play arms. Most ex…
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
Junyan Liu, Yunfan Li, Ruosong Wang +1
Existing metrics for reinforcement learning (RL) such as regret, PAC bounds, or uniform-PAC (Dann et al., 2017), typically evaluate the cumulative performance, while allowing the a…