activity
20242026
collaborators

7 papers

cs.LG2026

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

Mingyi Li, Taira Tsuchiya, Kenji Yamanishi

We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dep…

cs.LG2026

Data- and Variance-dependent Regret Bounds for Online Tabular MDPs

Mingyi Li, Taira Tsuchiya, Kenji Yamanishi

This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent re…

cs.LG2026

Bandit and Delayed Feedback in Online Structured Prediction

Yuki Shibukawa, Taira Tsuchiya, Shinsaku Sakaue +1

Online structured prediction is a task of sequentially predicting outputs with complex structures based on inputs and past observations, encompassing online classification. Recent…

eess.SY2025

Online Control of Linear Systems under Unbounded Noise

Kaito Ito, Taira Tsuchiya

This paper investigates the problem of controlling a linear system under possibly unbounded stochastic noise with unknown convex cost functions, known as an online control problem.…

cs.LG2025

Online Inverse Linear Optimization: Efficient Logarithmic-Regret Algorithm, Robustness to Suboptimality, and Lower Bound

Shinsaku Sakaue, Taira Tsuchiya, Han Bao +1

In online inverse linear optimization, a learner observes time-varying sets of feasible actions and an agent's optimal actions, selected by solving linear optimization over the fea…

cs.LG2025

Revisiting Online Learning Approach to Inverse Linear Optimization: A FenchelYoung Loss Perspective and Gap-Dependent Regret Analysis

Shinsaku Sakaue, Han Bao, Taira Tsuchiya

This paper revisits the online learning approach to inverse linear optimization studied by Bärmann et al. (2017), where the goal is to infer an unknown linear objective function o…