activity
20132026
most citedTowards Understanding the Importance of Shortcut Connections in Residual Networks

22 citations · 33 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2023

Bayesian Risk-Averse Q-Learning with Streaming Observations

Yuhao Wang, Enlu Zhou

We consider a robust reinforcement learning problem, where a learning agent learns from a simulated training environment. To account for the model mis-specification between this tr…

cs.LG2022

Noise Regularizes Over-parameterized Rank One Matrix Recovery, Provably

Tianyi Liu, Yan Li, Enlu Zhou +1

We investigate the role of noise in optimization algorithms for learning over-parameterized models. Specifically, we consider the recovery of a rank one matrix $Y^*\in R^{d\times d…

cs.LG20212 cited

Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix Factorization

Tianyi Liu, Yan Li, Song Wei +2

Numerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems. The theory behind such empirical observations, however, is still largely…

cs.LG201922 cited

Towards Understanding the Importance of Shortcut Connections in Residual Networks

Tianyi Liu, Minshuo Chen, Mo Zhou +3

Residual Network (ResNet) is undoubtedly a milestone in deep learning. ResNet is equipped with shortcut connections between layers, and exhibits efficient training using simple fir…

cs.LG2019

Towards Understanding the Importance of Noise in Training Neural Networks

Mo Zhou, Tianyi Liu, Yan Li +3

Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of neural networks. The theory behind, however, is still largel…

cs.LG2018

Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic Optimization

Tianyi Liu, Shiyang Li, Jianping Shi +2

Asynchronous momentum stochastic gradient descent algorithms (Async-MSGD) is one of the most popular algorithms in distributed machine learning. However, its convergence properties…