6 papers
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
Rei Higuchi, Ryotaro Kawata, Akifumi Wachi +3
Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depe…
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
Razvan-Andrei Lascu, Taiji Suzuki
We study the minimization of non-convex functionals over the Wasserstein space. While recent work has showed that perturbed Wasserstein gradient methods can avoid saddle points for…
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
Wei Huang, Andi Han, Mingyuan Bai +4
Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when da…
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
Ryoya Awano, Taiji Suzuki
Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman A…
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
Shokichi Takakura, Akifumi Wachi, Rei Higuchi +2
Aligning large language models (LLMs) to diverse human preferences is fundamentally challenging since criteria can often conflict with each other. Inference-time alignment methods…
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
Akifumi Wachi, Hirota Kinoshita, Shokichi Takakura +2
Reinforcement learning (RL) is a dominant paradigm for improving the reasoning abilities of large language models, yet its effectiveness varies across tasks and compute budgets. We…