5 citations · 6 across the 4 of their papers we have counts for
5 papers · 1 filter
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
Steven Kolawole, Lucio Dery, Jean-François Kagy +3
Structured pruning is a promising approach to create smaller, faster large language models. However, existing methods typically rely on computing the gradient via backward passes,…
Multitask Learning Can Improve Worst-Group Outcomes
Atharva Kulkarni, Lucio Dery, Amrith Setlur +3
In order to create machine learning systems that serve a variety of users well, it is vital to not only achieve high average performance but also ensure equitable outcomes across d…
Cross-Modal Fine-Tuning: Align then Refine
Junhong Shen, Liam Li, Lucio M. Dery +4
Fine-tuning large-scale pretrained models has led to tremendous progress in well-studied modalities such as vision and NLP. However, similar gains have not been observed in many ot…
Multi-step Planning for Automated Hyperparameter Optimization with OptFormer
Lucio M. Dery, Abram L. Friesen, Nando De Freitas +2
As machine learning permeates more industries and models become more expensive and time consuming to train, the need for efficient automated hyperparameter optimization (HPO) has n…
Auxiliary Task Update Decomposition: The Good, The Bad and The Neutral
Lucio M. Dery, Yann Dauphin, David Grangier
While deep learning has been very beneficial in data-rich settings, tasks with smaller training set often resort to pre-training or multitask learning to leverage data from other t…