activity
20162021
most citedFixup Initialization: Residual Learning Without Normalization

111 citations · 182 across the 3 of their papers we have counts for

collaborators

6 papers

cs.IR202145 cited

Context-Aware Legal Citation Recommendation using Deep Learning

Zihan Huang, Charles Low, Mengqiu Teng +4

Lawyers and judges spend a large amount of time researching the proper legal authority to cite while drafting decisions. In this paper, we develop a citation recommendation tool th…

cs.LG2021

One Backward from Ten Forward, Subsampling for Large-Scale Deep Learning

Chaosheng Dong, Xiaojie Jin, Weihao Gao +5

Deep learning models in large-scale machine learning systems are often continuously trained with enormous data from production environments. The sheer volume of streaming training…

cs.LG2019111 cited

Fixup Initialization: Residual Learning Without Normalization

Hongyi Zhang, Yann N. Dauphin, Tengyu Ma

Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate con…

math.OC201826 cited

R-SPIDER: A Fast Riemannian Stochastic Optimization Algorithm with Curvature Independent Rate

Jingzhao Zhang, Hongyi Zhang, Suvrit Sra

We study smooth stochastic optimization problems on Riemannian manifolds. Via adapting the recently proposed SPIDER algorithm \citep{fang2018spider} (a variance reduced stochastic…

math.OC2018

Towards Riemannian Accelerated Gradient Methods

Hongyi Zhang, Suvrit Sra

We propose a Riemannian version of Nesterov's Accelerated Gradient algorithm (RAGD), and show that for geodesically smooth and strongly convex problems, within a neighborhood of th…

math.OC2016

First-order Methods for Geodesically Convex Optimization

Hongyi Zhang, Suvrit Sra

Geodesic convexity generalizes the notion of (vector space) convexity to nonlinear metric spaces. But unlike convex optimization, geodesically convex (g-convex) optimization is muc…