22 citations · 23 across the 6 of their papers we have counts for
8 papers · 1 filter
Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry
Mo Zhou, Weihang Xu, Simon S. Du +1
Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, s…
Convergence Dynamics of Over-Parameterized Score Matching for a Single Gaussian
Yiran Zhang, Weihang Xu, Mo Zhou +2
Score matching has become a central training objective in modern generative modeling, particularly in diffusion models, where it is used to learn high-dimensional data distribution…
Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures
Mo Zhou, Weihang Xu, Maryam Fazel +1
Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradie…
How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
Mo Zhou, Rong Ge
The ability of learning useful features is one of the major advantages of neural networks. Although recent works show that neural network can operate in a neural tangent kernel (NT…
Depth Separation with Multilayer Mean-Field Networks
Yunwei Ren, Mo Zhou, Rong Ge
Depth separation -- why a deeper network is more powerful than a shallower one -- has been a major problem in deep learning theory. Previous results often focus on representation p…
A Local Convergence Theory for Mildly Over-Parameterized Two-Layer Neural Network
Mo Zhou, Rong Ge, Chi Jin
While over-parameterization is widely believed to be crucial for the success of optimization for the neural networks, most existing theories on over-parameterization do not fully e…