activity
20182023
most citedMaxUp: A Simple Way to Improve Generalization of Neural Network Training

34 citations · 102 across the 14 of their papers we have counts for

collaborators
Showing 2020 · cs.LGShow all

10 papers · 2 filters

cs.LG2020★ 3 cited

Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough

Mao Ye, Lemeng Wu, Qiang Liu

Despite the great success of deep learning, recent works show that large deep neural networks are often highly redundant and can be significantly reduced in size. However, the theo…

cs.LG2020

Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data

Mao Ye, Dhruv Choudhary, Jiecao Yu +6

Large scale deep learning provides a tremendous opportunity to improve the quality of content recommendation systems by employing both wider and deeper models, but this comes at gr…

cs.LG2020★ 5 cited

Go Wide, Then Narrow: Efficient Training of Deep Thin Networks

Denny Zhou, Mao Ye, Chen Chen +6

For deploying a deep learning model into production, it needs to be both accurate and compact to meet the latency and memory constraints. This usually results in a network that is…

cs.LG2020★ 8 cited

SAFER: A Structure-free Approach for Certified Robustness to Adversarial Word Substitutions

Mao Ye, Chengyue Gong, Qiang Liu

State-of-the-art NLP models can often be fooled by human-unaware transformations such as synonymous word substitution. For security reasons, it is of critical importance to develop…

cs.LG2020

Steepest Descent Neural Architecture Optimization: Escaping Local Optimum with Signed Neural Splitting

Lemeng Wu, Mao Ye, Qi Lei +2

Developing efficient and principled neural architecture optimization methods is a critical challenge of modern deep learning. Recently, Liu et al.[19] proposed a splitting steepest…

cs.LG2020

Good Subnetworks Provably Exist: Pruning via Greedy Forward Selection

Mao Ye, Chengyue Gong, Lizhen Nie +3

Recent empirical works show that large deep neural networks are often highly redundant and one can find much smaller subnetworks without a significant drop of accuracy. However, mo…