activity
20122026
most citedFine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

215 citations · 453 across the 21 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

cs.LG20227 cited

lo-fi: distributed fine-tuning without communication

Mitchell Wortsman, Suchin Gururangan, Shen Li +4

When fine-tuning large neural networks, it is common to use multiple nodes and to communicate gradients at each optimization step. By contrast, we investigate completely local fine…

cs.LG20212 cited

LCS: Learning Compressible Subspaces for Adaptive Network Compression at Inference Time

Elvis Nunez, Maxwell Horton, Anish Prabhu +3

When deploying deep learning models to a device, it is traditionally assumed that available computational resources (compute, memory, and power) remain static. However, real-world…

cs.LG2021

LLC: Accurate, Multi-purpose Learnt Low-dimensional Binary Codes

Aditya Kusupati, Matthew Wallingford, Vivek Ramanujan +6

Learning binary representations of instances and classes is a classical problem with several high potential applications. In modern settings, the compression of high-dimensional ne…

cs.LG2021

Learning Neural Network Subspaces

Mitchell Wortsman, Maxwell Horton, Carlos Guestrin +2

Recent observations have advanced our understanding of the neural network optimization landscape, revealing the existence of (1) paths of high accuracy containing diverse solutions…

cs.LG2020

Supermasks in Superposition

Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu +4

We present the Supermasks in Superposition (SupSup) model, capable of sequentially learning thousands of tasks without catastrophic forgetting. Our approach uses a randomly initial…