collaborators

15 papers

cs.AI2026

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

Sarvesh Gharat, Junpei Komiyama

Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. How…

cs.LG2026

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1

On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises…

cs.LG2026

Learning Optimal Dynamic Matching via Graph Neural Networks

Genta Okada, Shunya Noda, Junpei Komiyama +1

Dynamic matching markets require decisions about whom to match and when: matching now yields value but removes participants who may create better future opportunities. We develop a…

stat.ML2026

Fixed-Confidence Best-Arm Identification for Causal Mediation Analysis

Harsh Shrivastava, Yuta Kawakami, Junpei Komiyama +1

This paper studies the problem of identifying the treatment that maximizes the expected natural direct potential outcome (NDPO), which captures the potential outcome of an interven…

stat.ML2026

Replicability is Asymptotically Free in Multi-armed Bandits

Junpei Komiyama, Shinji Ito, Yuichi Yoshida +1

We consider a replicable stochastic multi-armed bandit algorithm that ensures, with high probability, that the algorithm's sequence of actions is not affected by the randomness inh…

cs.LG2026

Finite-Time Regret Analysis of Retry-Aware Bandits

Bingkui Tong, Junpei Komiyama, Soichiro Nishimori +1

We study a stochastic bandit algorithm motivated by retry-aware objectives that value the best outcome among multiple attempts, such as pass@ and max@. Given a posterior over…