11 citations · 44 across the 29 of their papers we have counts for
7 papers · 1 filter
High-dimensional density estimation with tensorizing flow
Yinuo Ren, Hongli Zhao, Yuehaw Khoo +1
We propose the tensorizing flow method for estimating high-dimensional probability density functions from the observed data. The method is based on tensor-train and flow-based gene…
Why self-attention is Natural for Sequence-to-Sequence Problems? A Perspective from Symmetries
Chao Ma, Lexing Ying
In this paper, we show that structures similar to self-attention are natural to learn many sequence-to-sequence problems from the perspective of symmetry. Inspired by language proc…
Bayesian regularization of empirical MDPs
Samarth Gupta, Daniel N. Hill, Lexing Ying +1
In most applications of model-based Markov decision processes, the parameters for the unknown underlying model are often estimated from the empirical data. Due to noise, the policy…
A Riemannian Mean Field Formulation for Two-layer Neural Networks with Batch Normalization
Chao Ma, Lexing Ying
The training dynamics of two-layer neural networks with batch normalization (BN) is studied. It is written as the training dynamics of a neural network without BN on a Riemannian m…
Combining resampling and reweighting for faithful stochastic optimization
Jing An, Lexing Ying
Many machine learning and data science tasks require solving non-convex optimization problems. When the loss function is a sum of multiple terms, a popular method is the stochastic…
How to Learn when Data Reacts to Your Model: Performative Gradient Descent
Zachary Izzo, Lexing Ying, James Zou
Performative distribution shift captures the setting where the choice of which ML model is deployed changes the data distribution. For example, a bank which uses the number of open…