163 citations · 220 across the 20 of their papers we have counts for
3 papers · 1 filter
TorchScale: Transformers at Scale
Shuming Ma, Hongyu Wang, Shaohan Huang +8
Large Transformers have achieved state-of-the-art performance across many tasks. Most open-source libraries on scaling Transformers focus on improving training or inference with be…
Foundation Transformers
Hongyu Wang, Shuming Ma, Shaohan Huang +12
A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…
Compression and Localization in Reinforcement Learning for ATARI Games
Joel Ruben Antony Moniz, Barun Patra, Sarthak Garg
Deep neural networks have become commonplace in the domain of reinforcement learning, but are often expensive in terms of the number of parameters needed. While compressing deep ne…