Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Distillation Scaling Laws
Dan Busbridge, Amitis Shidani, Floris Weers +3
We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate…
cs.LG2025
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
Samira Abnar, Harshay Shah, Dan Busbridge +3
Scaling the capacity of language models has consistently proven to be a reliable approach for improving performance and unlocking new capabilities. Capacity can be primarily define…
cs.LG2024
Poly-View Contrastive Learning
Amitis Shidani, Devon Hjelm, Jason Ramapuram +3
Contrastive learning typically matches pairs of related views among a number of unrelated negative views. Views can be generated (e.g. by augmentations) or be observed. We investig…