63 citations · 63 across the 1 of their papers we have counts for
3 papers
cs.CL2021★ 63 cited
BASE Layers: Simplifying Training of Large, Sparse Models
Mike Lewis, Shruti Bhosale, Tim Dettmers +2
We introduce a new balanced assignment of experts (BASE) layer for large language models that greatly simplifies existing high capacity sparse layers. Sparse layers can dramaticall…
cs.LG2019
Sparse Networks from Scratch: Faster Training without Losing Performance
Tim Dettmers, Luke Zettlemoyer
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense…
cs.CL2018
Jack the Reader - A Machine Reading Framework
Dirk Weissenborn, Pasquale Minervini, Tim Dettmers +8
Many Machine Reading and Natural Language Understanding tasks require reading supporting text in order to answer questions. For example, in Question Answering, the supporting text…