18 papers
Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization
Kenshi Abe, Mitsuki Sakamoto, Kaito Ariu +1
This paper proposes asymmetric perturbation, where only one player's payoff function is perturbed, for solving bilinear saddle-point optimization problems, commonly arising in mini…
Policy Testing in Markov Decision Processes
Kaito Ariu, Po-An Wang, Alexandre Proutiere +1
We study the policy testing problem in discounted Markov decision processes (MDPs) in the fixed-confidence setting under a generative model with static sampling. The goal is to dec…
Linear Convergence in Games with Delayed Feedback via Extra Prediction
Yuma Fujimoto, Kenshi Abe, Kaito Ariu
Feedback delays are inevitable in real-world multi-agent learning. They are known to severely degrade performance, and the convergence rate under delayed feedback is still unclear,…
Time-Varyingness in Auction Breaks Revenue Equivalence
Yuma Fujimoto, Kaito Ariu, Kenshi Abe
Auction is applied for trade with various mechanisms. A simple but practical question is which mechanism, typically first-price or second-price auctions, is preferred from the pers…
Consensus Group Relative Policy Optimization for Text Generation
Yuki Ichihara, Yuu Jinnai, Kaito Ariu +1
Many strong decoding methods for text generation follow a sample-and-rerank paradigm: they draw multiple candidates, score each under a utility (reward) function using consensus ac…
What you reward is what you learn: Comparing rewards for online speech policy optimization in public HRI
Sichao Song, Yuki Okafuji, Kaito Ariu +1
Designing policies that are both efficient and acceptable for conversational service robots in open and diverse environments is non-trivial. Unlike fixed, hand-tuned parameters, on…