118 citations · 127 across the 4 of their papers we have counts for
3 papers · 1 filter
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
Abhimanyu Rajeshkumar Bambhaniya, Amir Yazdanbakhsh, Suvinay Subramanian +4
N:M Structured sparsity has garnered significant interest as a result of relatively modest overhead and improved efficiency. Additionally, this form of sparsity holds considerable…
Scaling Laws for Sparsely-Connected Foundation Models
Elias Frantar, Carlos Riquelme, Neil Houlsby +2
We explore the impact of parameter sparsity on the scaling behavior of Transformers trained on massive datasets (i.e., "foundation models"), in both vision and language domains. In…
The Dormant Neuron Phenomenon in Deep Reinforcement Learning
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro +1
In this work we identify the dormant neuron phenomenon in deep reinforcement learning, where an agent's network suffers from an increasing number of inactive neurons, thereby affec…