117 citations · 145 across the 4 of their papers we have counts for
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2020
A Mean-field Analysis of Deep ResNet and Beyond: Towards Provable Optimization Via Overparameterization From Depth
Yiping Lu, Chao Ma, Yulong Lu +2
Training deep neural networks with stochastic gradient descent (SGD) can often achieve zero training loss on real-world tasks although the optimization landscape is known to be hig…
stat.ML2019★ 27 cited
Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
Bin Dong, Jikai Hou, Yiping Lu +1
Distillation is a method to transfer knowledge from one model to another and often achieves higher accuracy with the same capacity. In this paper, we aim to provide a theoretical u…
stat.ML2019
You Only Propagate Once: Accelerating Adversarial Training via Maximal Principle
Dinghuai Zhang, Tianyuan Zhang, Yiping Lu +2
Deep learning achieves state-of-the-art results in many tasks in computer vision and natural language processing. However, recent works have shown that deep networks can be vulnera…