1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CL2025
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
cs.IR2023
Exploration of Unranked Items in Safe Online Learning to Re-Rank
Hiroaki Shiino, Kaito Ariu, Kenshi Abe +1
Bandit algorithms for online learning to rank (OLTR) problems often aim to maximize long-term revenue by utilizing user feedback. From a practical point of view, however, such algo…
cs.GT2023★ 1 cited
Learning in Multi-Memory Games Triggers Complex Dynamics Diverging from Nash Equilibrium
Yuma Fujimoto, Kaito Ariu, Kenshi Abe
Repeated games consider a situation where multiple agents are motivated by their independent rewards throughout learning. In general, the dynamics of their learning become complex.…