4 papers
Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
Diffusion Language Models (DLMs) have recently become increasingly competitive with autoregressive (AR) models, and even outperform them on certain tasks. Unlike AR models, DLMs pr…
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises…
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
Theoretical Guarantees for Minimum Bayes Risk Decoding
Yuki Ichihara, Yuu Jinnai, Kaito Ariu +2
Minimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution. While prior work has shown the effectiven…