3 papers
cs.GT2026
Policy Iteration for Two-Player General-Sum Stochastic Stackelberg Games
Mikoto Kudo, Youhei Akimoto
We address two-player general-sum stochastic Stackelberg games (SSGs), where the leader's policy is optimized considering the best-response follower whose policy is optimal for its…
stat.ML2024
Harnessing the Power of Vicinity-Informed Analysis for Classification under Covariate Shift
Mitsuhiro Fujikawa, Yohei Akimoto, Jun Sakuma +1
Transfer learning enhances prediction accuracy on a target distribution by leveraging data from a source distribution, demonstrating significant benefits in various applications. T…
cs.LG2024
Stepwise Alignment for Constrained Language Model Policy Optimization
Akifumi Wachi, Thien Q. Tran, Rei Sato +2
Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment…