4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.DC2022★ 4 cited
Distributed SLIDE: Enabling Training Large Neural Networks on Low Bandwidth and Simple CPU-Clusters via Model Parallelism and Sparsity
Minghao Yan, Nicholas Meisburger, Tharun Medini +1
More than 70% of cloud computing is paid for but sits idle. A large fraction of these idle compute are cheap CPUs with few cores that are not utilized during the less busy hours. T…
cs.LG2021★ 1 cited
PairConnect: A Compute-Efficient MLP Alternative to Attention
Zhaozhuo Xu, Minghao Yan, Junyan Zhang +1
Transformer models have demonstrated superior performance in natural language processing. The dot product self-attention in Transformer allows us to model interactions between word…