Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
Rei Higuchi, Ryotaro Kawata, Akifumi Wachi +3
Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depe…
stat.ML2026
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
Shokichi Takakura, Akifumi Wachi, Rei Higuchi +2
Aligning large language models (LLMs) to diverse human preferences is fundamentally challenging since criteria can often conflict with each other. Inference-time alignment methods…