145 citations · 176 across the 7 of their papers we have counts for
4 papers · 1 filter
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
Aleksandar Botev, Soham De, Samuel L Smith +59
We introduce RecurrentGemma, a family of open language models which uses Google's novel Griffin architecture. Griffin combines linear recurrences with local attention to achieve ex…
Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation
Bobby He, James Martens, Guodong Zhang +4
Skip connections and normalisation layers form two standard architectural components that are ubiquitous for the training of Deep Neural Networks (DNNs), but whose precise roles ar…
Spatial Functa: Scaling Functa to ImageNet Classification and Generation
Matthias Bauer, Emilien Dupont, Andy Brock +3
Neural fields, also known as implicit neural representations, have emerged as a powerful means to represent complex signals of various modalities. Based on this Dupont et al. (2022…
TF-Replicator: Distributed Machine Learning for Researchers
Peter Buchlovsky, David Budden, Dominik Grewe +9
We describe TF-Replicator, a framework for distributed machine learning designed for DeepMind researchers and implemented as an abstraction over TensorFlow. TF-Replicator simplifie…