3 citations · 3 across the 1 of their papers we have counts for
3 papers
Scaling Laws for Sparsely-Connected Foundation Models
Elias Frantar, Carlos Riquelme, Neil Houlsby +2
We explore the impact of parameter sparsity on the scaling behavior of Transformers trained on massive datasets (i.e., "foundation models"), in both vision and language domains. In…
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
The Dormant Neuron Phenomenon in Deep Reinforcement Learning
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro +1
In this work we identify the dormant neuron phenomenon in deep reinforcement learning, where an agent's network suffers from an increasing number of inactive neurons, thereby affec…