3 papers
cs.LG2026
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dep…
cs.LG2026
Data- and Variance-dependent Regret Bounds for Online Tabular MDPs
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent re…
cs.LG2025
Bandit and Delayed Feedback in Online Structured Prediction
Yuki Shibukawa, Taira Tsuchiya, Shinsaku Sakaue +1
Online structured prediction is a task of sequentially predicting outputs with complex structures based on inputs and past observations, encompassing online classification. Recent…