4 papers
When Does Preconditioning Help or Hurt Generalization?
Shun-ichi Amari, Jimmy Ba, Roger Grosse +5
While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a…
Scalable Gradients for Stochastic Differential Equations
Xuechen Li, Ting-Kam Leonard Wong, Ricky T. Q. Chen +1
The adjoint sensitivity method scalably computes gradients of solutions to ordinary differential equations. We generalize this method to stochastic differential equations, allowing…
Stochastic Runge-Kutta Accelerates Langevin Monte Carlo and Beyond
Xuechen Li, Denny Wu, Lester Mackey +1
Sampling with Markov chain Monte Carlo methods often amounts to discretizing some continuous-time dynamics with numerical integration. In this paper, we establish the convergence r…
Isolating Sources of Disentanglement in Variational Autoencoders
Ricky T. Q. Chen, Xuechen Li, Roger Grosse +1
We decompose the evidence lower bound to show the existence of a term measuring the total correlation between latent variables. We use this to motivate our -TCVAE (Total Correla…