activity
20242026
collaborators

15 papers

cs.LG2026

Combinatorial Allocation Bandits with Nonlinear Arm Utility

Yuki Shibukawa, Koichi Tanaka, Yuta Saito +1

A matching platform is a system that matches participants of different types, such as companies and job-seekers. In such a platform, maximizing matches may concentrate assignments…

cs.LG2026

A Perturbation Approach to Unconstrained Linear Bandits

Andrew Jacobsen, Dorian Baudry, Shinji Ito +1

We revisit the standard perturbation-based approach of Abernethy et al. (2008) in the context of unconstrained Bandit Linear Optimization (uBLO). We show the surprising result that…

cs.GT2026

Scale-Invariant Fast Convergence in Games

Taira Tsuchiya, Haipeng Luo, Shinji Ito

Scale-invariance in games has recently emerged as a widely valued desirable property. Yet, almost all fast convergence guarantees in learning in games require prior knowledge of th…

cs.LG2026

Adversarial Learning in Games with Bandit Feedback: Logarithmic Pure-Strategy Maximin Regret

Shinji Ito, Haipeng Luo, Arnab Maiti +2

Learning to play zero-sum games is a fundamental problem in game theory and machine learning. While significant progress has been made in minimizing external regret in the self-pla…

cs.LG2025

Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback

Shinji Ito, Kevin Jamieson, Haipeng Luo +2

We study online learning in finite-horizon episodic Markov decision processes (MDPs) under the challenging aggregate bandit feedback model, where the learner observes only the cumu…

cs.LG2025

Reinforcement Learning from Adversarial Preferences in Tabular MDPs

Taira Tsuchiya, Shinji Ito, Haipeng Luo

We introduce a new framework of episodic tabular Markov decision processes (MDPs) with adversarial preferences, which we refer to as preference-based MDPs (PbMDPs). Unlike standard…