2 papers
cs.LG2026
A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets
Akihiro Kubo, Kosuke Nakanishi, Shin Ishii
Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonlinear scalarization the uniqu…
cs.LG2025
Double Horizon Model-Based Policy Optimization
Akihiro Kubo, Paavo Parmas, Shin Ishii
Model-based reinforcement learning (MBRL) reduces the cost of real-environment sampling by generating synthetic trajectories (called rollouts) from a learned dynamics model. Howeve…