15 papers
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
Sarvesh Gharat, Junpei Komiyama
Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. How…
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises…
Learning Optimal Dynamic Matching via Graph Neural Networks
Genta Okada, Shunya Noda, Junpei Komiyama +1
Dynamic matching markets require decisions about whom to match and when: matching now yields value but removes participants who may create better future opportunities. We develop a…
Fixed-Confidence Best-Arm Identification for Causal Mediation Analysis
Harsh Shrivastava, Yuta Kawakami, Junpei Komiyama +1
This paper studies the problem of identifying the treatment that maximizes the expected natural direct potential outcome (NDPO), which captures the potential outcome of an interven…
Replicability is Asymptotically Free in Multi-armed Bandits
Junpei Komiyama, Shinji Ito, Yuichi Yoshida +1
We consider a replicable stochastic multi-armed bandit algorithm that ensures, with high probability, that the algorithm's sequence of actions is not affected by the randomness inh…
Finite-Time Regret Analysis of Retry-Aware Bandits
Bingkui Tong, Junpei Komiyama, Soichiro Nishimori +1
We study a stochastic bandit algorithm motivated by retry-aware objectives that value the best outcome among multiple attempts, such as pass@ and max@. Given a posterior over…