128 citations · 136 across the 8 of their papers we have counts for
6 papers · 1 filter
Distillation Scaling Laws
Dan Busbridge, Amitis Shidani, Floris Weers +3
We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate…
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
Samira Abnar, Harshay Shah, Dan Busbridge +3
Scaling the capacity of language models has consistently proven to be a reliable approach for improving performance and unlocking new capabilities. Capacity can be primarily define…
Elastic Weight Consolidation Improves the Robustness of Self-Supervised Learning Methods under Transfer
Andrius Ovsianas, Jason Ramapuram, Dan Busbridge +2
Self-supervised representation learning (SSL) methods provide an effective label-free initial condition for fine-tuning downstream tasks. However, in numerous realistic scenarios,…
Evaluating the fairness of fine-tuning strategies in self-supervised learning
Jason Ramapuram, Dan Busbridge, Russ Webb
In this work we examine how fine-tuning impacts the fairness of contrastive Self-Supervised Learning (SSL) models. Our findings indicate that Batch Normalization (BN) statistics pl…
Neural Temporal Point Processes For Modelling Electronic Health Records
Joseph Enguehard, Dan Busbridge, Adam Bozson +2
The modelling of Electronic Health Records (EHRs) has the potential to drive more efficient allocation of healthcare resources, enabling early intervention strategies and advancing…
Relational Graph Attention Networks
Dan Busbridge, Dane Sherburn, Pietro Cavallo +1
We investigate Relational Graph Attention Networks, a class of models that extends non-relational graph attention mechanisms to incorporate relational information, opening up these…