4 papers
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
Rei Higuchi, Ryotaro Kawata, Akifumi Wachi +3
Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depe…
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
Shokichi Takakura, Akifumi Wachi, Rei Higuchi +2
Aligning large language models (LLMs) to diverse human preferences is fundamentally challenging since criteria can often conflict with each other. Inference-time alignment methods…
Path Learning with Trajectory Advantage Regression
Kohei Miyaguchi
In this paper, we propose trajectory advantage regression, a method of offline path learning and path attribution based on reinforcement learning. The proposed method can be used t…
A Provable Approach for End-to-End Safe Reinforcement Learning
Akifumi Wachi, Kohei Miyaguchi, Takumi Tanabe +2
A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing sa…