50 citations · 97 across the 9 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022
Transformer with Memory Replay
Rui Liu, Barzan Mozafari
Transformers achieve state-of-the-art performance for natural language processing tasks by pre-training on large-scale text corpora. They are extremely compute-intensive and have v…
cs.LG2022★ 11 cited
Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers
Rui Liu, Young Jin Kim, Alexandre Muzio +1
Sparsely activated transformers, such as Mixture of Experts (MoE), have received great interest due to their outrageous scaling capability which enables dramatical increases in mod…