1k citations · 2.2k across the 20 of their papers we have counts for
3 papers · 1 filter
Deconstructing the Structure of Sparse Neural Networks
Maxwell Van Gelder, Mitchell Wortsman, Kiana Ehsani
Although sparse neural networks have been studied extensively, the focus has been primarily on accuracy. In this work, we focus instead on network structure, and analyze three popu…
Supermasks in Superposition
Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu +4
We present the Supermasks in Superposition (SupSup) model, capable of sequentially learning thousands of tasks without catastrophic forgetting. Our approach uses a randomly initial…
Soft Threshold Weight Reparameterization for Learnable Sparsity
Aditya Kusupati, Vivek Ramanujan, Raghav Somani +4
Sparsity in Deep Neural Networks (DNNs) is studied extensively with the focus of maximizing prediction accuracy given an overall parameter budget. Existing methods rely on uniform…