3 papers
cs.CL2025
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
cs.GT2025
On the Power of Perturbation under Sampling in Solving Extensive-Form Games
Wataru Masaka, Mitsuki Sakamoto, Kenshi Abe +3
We investigate how perturbation does and does not improve the Follow-the-Regularized-Leader (FTRL) algorithm in solving imperfect-information extensive-form games under sampling, w…
cs.GT2024
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games
Kenshi Abe, Mitsuki Sakamoto, Kaito Ariu +1
This paper presents a payoff perturbation technique, introducing a strong convexity to players' payoff functions in games. This technique is specifically designed for first-order m…