4 citations · 8 across the 7 of their papers we have counts for
7 papers
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
Scalable and Provably Fair Exposure Control for Large-Scale Recommender Systems
Riku Togashi, Kenshi Abe, Yuta Saito
Typical recommendation and ranking methods aim to optimize the satisfaction of users, but they are often oblivious to their impact on the items (e.g., products, jobs, news, video)…
Learning Fair Division from Bandit Feedback
Hakuei Yamada, Junpei Komiyama, Kenshi Abe +1
This work addresses learning online fair division under uncertainty, where a central planner sequentially allocates items without precise knowledge of agents' values or utilities.…
Why Guided Dialog Policy Learning performs well? Understanding the role of adversarial learning and its alternative
Sho Shimoyama, Tetsuro Morimura, Kenshi Abe +6
Dialog policies, which determine a system's action based on the current state at each dialog turn, are crucial to the success of the dialog. In recent years, reinforcement learning…
Exploration of Unranked Items in Safe Online Learning to Re-Rank
Hiroaki Shiino, Kaito Ariu, Kenshi Abe +1
Bandit algorithms for online learning to rank (OLTR) problems often aim to maximize long-term revenue by utilizing user feedback. From a practical point of view, however, such algo…
Learning in Multi-Memory Games Triggers Complex Dynamics Diverging from Nash Equilibrium
Yuma Fujimoto, Kaito Ariu, Kenshi Abe
Repeated games consider a situation where multiple agents are motivated by their independent rewards throughout learning. In general, the dynamics of their learning become complex.…