2 papers
cs.LG2024
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
Aleksandar Botev, Soham De, Samuel L Smith +59
We introduce RecurrentGemma, a family of open language models which uses Google's novel Griffin architecture. Griffin combines linear recurrences with local attention to achieve ex…
cs.LG2024
Scorch: A Library for Sparse Deep Learning
Bobby Yan, Alexander J. Root, Trevor Gale +2
The rapid growth in the size of deep learning models strains the capabilities of traditional dense computation paradigms. Leveraging sparse computation has become increasingly popu…