Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dep…
cs.LG2026
Data- and Variance-dependent Regret Bounds for Online Tabular MDPs
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent re…
cs.LG2025
Modified K-means Algorithm with Local Optimality Guarantees
Mingyi Li, Michael R. Metel, Akiko Takeda
The K-means algorithm is one of the most widely studied clustering algorithms in machine learning. While extensive research has focused on its ability to achieve a globally optimal…