6 papers
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Maciej ChrabÄ szcz, Aleksander Szymczyk, Marcin Sendera +2
Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's f…
From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training
Julius Berner, Lorenz Richter, Marcin Sendera +2
We study the problem of training neural stochastic differential equations, or diffusion models, to sample from a Boltzmann distribution without access to target samples. Existing m…
Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models
Siddarth Venkatraman, Mohsin Hasan, Minsu Kim +5
Any well-behaved generative model over a variable can be expressed as a deterministic transformation of an exogenous ('outsourced') Gaussian noise variable $\mathbf{z}…
Solving Bayesian inverse problems with diffusion priors and off-policy RL
Luca Scimeca, Siddarth Venkatraman, Moksh Jain +14
This paper presents a practical application of Relative Trajectory Balance (RTB), a recently introduced off-policy reinforcement learning (RL) objective that can asymptotically sol…
Amortizing intractable inference in diffusion models for vision, language, and control
Siddarth Venkatraman, Moksh Jain, Luca Scimeca +12
Diffusion models have emerged as effective distribution estimators in vision, language, and reinforcement learning, but their use as priors in downstream tasks poses an intractable…
Improved off-policy training of diffusion samplers
Marcin Sendera, Minsu Kim, Sarthak Mittal +6
We study the problem of training diffusion models to sample from a distribution with a given unnormalized density or energy function. We benchmark several diffusion-structured infe…