9 papers
Adaptive Calibration in Non-Stationary Environments
Junyan Liu, Haipeng Luo, Lillian J. Ratliff
Making calibrated online predictions is a central challenge in modern AI systems. Much of the existing literature focuses on fully adversarial environments where outcomes may be ar…
Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning
Max Horwitz, Jake Gonzales, Eric Mazumdar +1
A growing line of work reframes preference-based fine-tuning of large language models game-theoretically: Nash Learning from Human Feedback (NLHF) recasts the problem as a zero-sum…
Safe Probabilistic Planning for Human-Robot Interaction using Conformal Risk Control
Jake Gonzales, Kazuki Mizuta, Karen Leung +1
In this paper, we present a novel probabilistic safe control framework for human-robot interaction that combines control barrier functions (CBFs) with conformal risk control to pro…
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
Shirui Chen, Cole Harrison, Ying-Chun Lee +6
General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.…
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
Junyan Liu, Haipeng Luo, Zihan Zhang +1
We study online learning in two-player uninformed Markov games, where the opponent's actions and policies are unobserved. In this setting, Tian et al. (2021) show that achieving no…
Improved Regret and Contextual Linear Extension for Pandora's Box and Prophet Inequality
Junyan Liu, Ziyun Chen, Kun Wang +2
We study the Pandora's Box problem in an online learning setting with semi-bandit feedback. In each round, the learner sequentially pays to open up to boxes with unknown reward…