Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
Rei Higuchi, Ryotaro Kawata, Akifumi Wachi +3
Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depe…
stat.ML2026
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
Ryotaro Kawata, Taiji Suzuki
Transformers excel through content-addressable retrieval and the ability to exploit contexts of, in principle, unbounded length. We recast associative memory at the level of probab…