368 citations · 921 across the 30 of their papers we have counts for
Showing 2018Show all
2 papers · 1 filter
cs.LG2018
WEST: Word Encoded Sequence Transducers
Ehsan Variani, Ananda Theertha Suresh, Mitchel Weintraub
Most of the parameters in large vocabulary models are used in embedding layer to map categorical features to vectors and in softmax layer for classification weights. This is a bott…
stat.ML2018
cpSGD: Communication-efficient and differentially-private distributed SGD
Naman Agarwal, Ananda Theertha Suresh, Felix Yu +2
Distributed stochastic gradient descent is an important subroutine in distributed learning. A setting of particular interest is when the clients are mobile devices, where two impor…