Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling
arXiv:2609.05727
Abstract
We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is , where is the reward, the inverse temperature, and denotes the pretrained model's terminal density for fine-tuning or the constant for sampling. We shift the paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities. Under compatible smooth-realization assumptions, canonical velocities form a manifold diffeomorphic to the density manifold. Transporting the Fisher-Rao metric and mixture connection to this manifold, we show that the reverse-KL Hessian equals the metric, so the Newton direction coincides with the negative Fisher-Rao gradient. At terminal density , each stage takes a tangential step generated by the regularized reward , followed by terminal-density-preserving canonicalization. This canonical retraction yields an exact finite-stepsize density characterization. For the ideal iteration, we prove strict reverse-KL descent away from the target for , global convergence under mild conditions, and local quadratic convergence for full steps (). Covariance and gradient forms, each with forward or reverse regression-pair constructions, yield sample-wise tangential-update losses with the same population minimizer, without importance sampling or full-trajectory backpropagation. We develop approximate updates and define critical-point consistency as vanishing tangential displacement if and only if . We recover representative methods as exact realizations, critical-point-consistent approximations, or objective-altering variants, enabling modular algorithm design. Our work advances the theory and algorithms of reinforcement learning for generative models.