activity
20242026
collaborators
Showing 2025Show all

5 papers · 1 filter

cs.LG2025

Latent Variable Modeling for Robust Causal Effect Estimation

Tetsuro Morimura, Tatsushi Oka, Yugo Suzuki +1

Latent variable models provide a powerful framework for incorporating and inferring unobserved factors in observational data. In causal inference, they help account for hidden fact…

cs.CL2025

Theoretical Guarantees for Minimum Bayes Risk Decoding

Yuki Ichihara, Yuu Jinnai, Kaito Ariu +2

Minimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution. While prior work has shown the effectiven…

cs.LG2025

Return-Aligned Decision Transformer

Tsunehiko Tanaka, Kenshi Abe, Kaito Ariu +2

Traditional approaches in offline reinforcement learning aim to learn the optimal policy that maximizes the cumulative reward, also known as return. It is increasingly important to…

cs.CL2025

Evaluation of Best-of-N Sampling Strategies for Language Model Alignment

Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4

Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…

cs.CL2025

Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Yuu Jinnai, Tetsuro Morimura, Kaito Ariu +1

Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) to human preferences at the time of decoding. BoN…