3 papers
cs.CL2025
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
cs.LG2020
A Practical Guide of Off-Policy Evaluation for Bandit Problems
Masahiro Kato, Kenshi Abe, Kaito Ariu +1
Off-policy evaluation (OPE) is the problem of estimating the value of a target policy from samples obtained via different policies. Recently, applying OPE methods for bandit proble…
stat.ML2020
Regret in Online Recommendation Systems
Kaito Ariu, Narae Ryu, Se-Young Yun +1
This paper proposes a theoretical analysis of recommendation systems in an online setting, where items are sequentially recommended to users over time. In each round, a user, rando…