151 citations · 389 across the 13 of their papers we have counts for
5 papers · 1 filter
Revisiting BFloat16 Training
Pedram Zamirai, Jian Zhang, Christopher R. Aberger +1
State-of-the-art generic low-precision training algorithms use a mix of 16-bit and 32-bit precision, creating the folklore that 16-bit hardware compute units alone are not enough t…
Neural Manifold Ordinary Differential Equations
Aaron Lou, Derek Lim, Isay Katsman +4
To better conform to data geometry, recent deep generative modelling techniques adapt Euclidean constructions to non-Euclidean spaces. In this paper, we study normalizing flows on…
MixML: A Unified Analysis of Weakly Consistent Parallel Learning
Yucheng Lu, Jack Nash, Christopher De Sa
Parallelism is a ubiquitous method for accelerating machine learning algorithms. However, theoretical analysis of parallel learning is usually done in an algorithm- and protocol-sp…
Optimizing JPEG Quantization for Classification Networks
Zhijing Li, Christopher De Sa, Adrian Sampson
Deep learning for computer vision depends on lossy image compression: it reduces the storage required for training and test data and lowers transfer costs in deployment. Mainstream…
Moniqua: Modulo Quantized Communication in Decentralized SGD
Yucheng Lu, Christopher De Sa
Running Stochastic Gradient Descent (SGD) in a decentralized fashion has shown promising results. In this paper we propose Moniqua, a technique that allows decentralized SGD to use…