7 papers
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dep…
Data- and Variance-dependent Regret Bounds for Online Tabular MDPs
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent re…
Bandit and Delayed Feedback in Online Structured Prediction
Yuki Shibukawa, Taira Tsuchiya, Shinsaku Sakaue +1
Online structured prediction is a task of sequentially predicting outputs with complex structures based on inputs and past observations, encompassing online classification. Recent…
Online Control of Linear Systems under Unbounded Noise
Kaito Ito, Taira Tsuchiya
This paper investigates the problem of controlling a linear system under possibly unbounded stochastic noise with unknown convex cost functions, known as an online control problem.…
Online Inverse Linear Optimization: Efficient Logarithmic-Regret Algorithm, Robustness to Suboptimality, and Lower Bound
Shinsaku Sakaue, Taira Tsuchiya, Han Bao +1
In online inverse linear optimization, a learner observes time-varying sets of feasible actions and an agent's optimal actions, selected by solving linear optimization over the fea…
Revisiting Online Learning Approach to Inverse Linear Optimization: A FenchelYoung Loss Perspective and Gap-Dependent Regret Analysis
Shinsaku Sakaue, Han Bao, Taira Tsuchiya
This paper revisits the online learning approach to inverse linear optimization studied by Bärmann et al. (2017), where the goal is to infer an unknown linear objective function o…