3 papers
stat.ML2026
Policy Testing in Markov Decision Processes
Kaito Ariu, Po-An Wang, Alexandre Proutiere +1
We study the policy testing problem in discounted Markov decision processes (MDPs) in the fixed-confidence setting under a generative model with static sampling. The goal is to dec…
cs.CL2025
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +4
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. Bo…
cs.LG2024
Mean-Variance Efficient Reinforcement Learning with Applications to Dynamic Financial Investment
Masahiro Kato, Kei Nakagawa, Kenshi Abe +2
This study investigates the mean-variance (MV) trade-off in reinforcement learning (RL), an instance of the sequential decision-making under uncertainty. Our objective is to obtain…