1 paper · 1 filter
Ankur Moitra, Andrej Risteski, Dhruv Rohatgi
Inference-time reward alignment asks how to turn a pre-trained diffusion model with base law p into a sampler that favors a reward r while remaining close to p. Since there i…