activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +3

Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available. However, such a highly reliable reward model is no…

cs.LG2025

Latent Variable Modeling for Robust Causal Effect Estimation

Tetsuro Morimura, Tatsushi Oka, Yugo Suzuki +1

Latent variable models provide a powerful framework for incorporating and inferring unobserved factors in observational data. In causal inference, they help account for hidden fact…

cs.LG2025

Return-Aligned Decision Transformer

Tsunehiko Tanaka, Kenshi Abe, Kaito Ariu +2

Traditional approaches in offline reinforcement learning aim to learn the optimal policy that maximizes the cumulative reward, also known as return. It is increasingly important to…

cs.LG2024

Filtered Direct Preference Optimization

Tetsuro Morimura, Mitsuki Sakamoto, Yuu Jinnai +2

Reinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences. While the significance of dataset quality is generally re…

cs.LG2024

Mean-Variance Efficient Reinforcement Learning with Applications to Dynamic Financial Investment

Masahiro Kato, Kei Nakagawa, Kenshi Abe +2

This study investigates the mean-variance (MV) trade-off in reinforcement learning (RL), an instance of the sequential decision-making under uncertainty. Our objective is to obtain…

cs.LG2024

Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes

Tetsuro Morimura, Kazuhiro Ota, Kenshi Abe +1

Policy gradient (PG) is a reinforcement learning (RL) approach that optimizes a parameterized policy model for an expected return using gradient ascent. While PG can work well even…