activity
20182022
most citedMixML: A Unified Analysis of Weakly Consistent Parallel Learning

4 citations · 11 across the 3 of their papers we have counts for

collaborators

7 papers

cs.LG20224 cited

Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam

Yucheng Lu, Conglong Li, Minjia Zhang +2

1-bit gradient compression and local steps are two representative techniques that enable drastic communication reduction in distributed SGD. Their benefits, however, remain an open…

cs.LG2021

Iterated learning for emergent systematicity in VQA

Ankit Vani, Max Schwarzer, Yuchen Lu +2

Although neural module networks have an architectural bias towards compositionality, they require gold standard layouts to generalize systematically in practice. When instead learn…

cs.LG20213 cited

Learning Task Decomposition with Ordered Memory Policy Network

Yuchen Lu, Yikang Shen, Siyuan Zhou +3

Many complex real-world tasks are composed of several levels of sub-tasks. Humans leverage these hierarchical structures to accelerate the learning process and achieve better gener…

cs.LG2021

Variance Reduced Training with Stratified Sampling for Forecasting Models

Yucheng Lu, Youngsuk Park, Lifan Chen +3

In large-scale time series forecasting, one often encounters the situation where the temporal patterns of time series, while drifting over time, differ from one another in the same…

cs.LG20204 cited

MixML: A Unified Analysis of Weakly Consistent Parallel Learning

Yucheng Lu, Jack Nash, Christopher De Sa

Parallelism is a ubiquitous method for accelerating machine learning algorithms. However, theoretical analysis of parallel learning is usually done in an algorithm- and protocol-sp…

cs.LG2020

Moniqua: Modulo Quantized Communication in Decentralized SGD

Yucheng Lu, Christopher De Sa

Running Stochastic Gradient Descent (SGD) in a decentralized fashion has shown promising results. In this paper we propose Moniqua, a technique that allows decentralized SGD to use…