activity
20192026
most citedTowards Understanding the Importance of Shortcut Connections in Residual Networks

22 citations · 23 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

Mo Zhou, Weihang Xu, Simon S. Du +1

Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, s…

cs.LG2025

Convergence Dynamics of Over-Parameterized Score Matching for a Single Gaussian

Yiran Zhang, Weihang Xu, Mo Zhou +2

Score matching has become a central training objective in modern generative modeling, particularly in diffusion models, where it is used to learn high-dimensional data distribution…

cs.LG2025

Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

Mo Zhou, Weihang Xu, Maryam Fazel +1

Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradie…

cs.LG2024

How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks

Mo Zhou, Rong Ge

The ability of learning useful features is one of the major advantages of neural networks. Although recent works show that neural network can operate in a neural tangent kernel (NT…

cs.LG2023

Depth Separation with Multilayer Mean-Field Networks

Yunwei Ren, Mo Zhou, Rong Ge

Depth separation -- why a deeper network is more powerful than a shallower one -- has been a major problem in deep learning theory. Previous results often focus on representation p…

cs.LG2021

A Local Convergence Theory for Mildly Over-Parameterized Two-Layer Neural Network

Mo Zhou, Rong Ge, Chi Jin

While over-parameterization is widely believed to be crucial for the success of optimization for the neural networks, most existing theories on over-parameterization do not fully e…