6 papers
Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization
Kenshi Abe, Mitsuki Sakamoto, Kaito Ariu +1
This paper proposes asymmetric perturbation, where only one player's payoff function is perturbed, for solving bilinear saddle-point optimization problems, commonly arising in mini…
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +3
Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available. However, such a highly reliable reward model is no…
On the Power of Perturbation under Sampling in Solving Extensive-Form Games
Wataru Masaka, Mitsuki Sakamoto, Kenshi Abe +3
We investigate how perturbation does and does not improve the Follow-the-Regularized-Leader (FTRL) algorithm in solving imperfect-information extensive-form games under sampling, w…
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games
Kenshi Abe, Mitsuki Sakamoto, Kaito Ariu +1
This paper presents a payoff perturbation technique, introducing a strong convexity to players' payoff functions in games. This technique is specifically designed for first-order m…
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
Filtered Direct Preference Optimization
Tetsuro Morimura, Mitsuki Sakamoto, Yuu Jinnai +2
Reinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences. While the significance of dataset quality is generally re…