4 papers
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises…
Reliable Chain-of-Thought via Prefix Consistency
Naoto Iwase, Yuki Ichihara, Mohammad Atif Quamar +1
Large Language Models often improve accuracy on reasoning tasks by sampling multiple Chain-of-Thought (CoT) traces and aggregating them with majority voting (MV), a test-time techn…
CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency
Hirofumi Ota, Naoto Iwase, Yuki Ichihara +2
Large language models often improve reasoning by sampling multiple outputs and aggregating their final answers, but precise and efficient control of error levels remains a challeng…
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
Naoto Iwase, Hiroki Okuyama, Junichiro Iwasawa
Large language models (LLMs) show increasing promise in medical applications, but their ability to detect and correct errors in clinical texts -- a prerequisite for safe deployment…