activity
20212026
most citedShinRL: A Library for Evaluating RL Algorithms from Theoretical and Practical Perspectives

1 citations · 1 across the 13 of their papers we have counts for

collaborators

14 papers

cs.LG2026

Randomized Exploration for Linear Bandits via Absolute Perturbations

Toshinori Kitamura, Shuai Liu, Csaba Szepesvári

In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson…

cs.LG2026

Offline-to-Online Learning in Linear Bandits

Kushagra Chandak, Toshinori Kitamura, Xiaoqi Tan

We study online learning with an additional offline dataset in the stochastic linear bandit setting. Although this problem arises frequently in practice, the offline-to-online trad…

cs.LG2026

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada +4

In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce un…

math.OC2026

Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions

Toshinori Kitamura, Arnob Ghosh, Alex Ayoub +2

Projected subgradient descent (PSD) has gained popularity for solving robust Markov decision processes (RMDPs) because it applies to a broader class of uncertainty sets than tradit…

cs.RO2025

A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics

Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa +8

Recent Foundation Model-enabled robotics (FMRs) display greatly improved general-purpose skills, enabling more adaptable automation than conventional robotics. Their ability to han…

cs.LG2025

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

Toshinori Kitamura, Arnob Ghosh, Tadashi Kozuno +5

We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward…