243 citations · 402 across the 12 of their papers we have counts for
30 papers
Unified Scaling Laws for Routed Language Models
Aidan Clark, Diego de las Casas, Aurelia Guy +23
The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: archite…
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…
Prune Responsibly
Michela Paganini
Irrespective of the specific definition of fairness in a machine learning application, pruning the underlying model affects it. We investigate and document the emergence and exacer…
Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding
Michela Paganini, Jessica Zosa Forde
The observation of sparse trainable sub-networks within over-parametrized networks - also known as Lottery Tickets (LTs) - has prompted inquiries around their trainability, scaling…
dagger: A Python Framework for Reproducible Machine Learning Experiment Orchestration
Michela Paganini, Jessica Zosa Forde
Many research directions in machine learning, particularly in deep learning, involve complex, multi-stage experiments, commonly involving state-mutating operations acting on models…
Streamlining Tensor and Network Pruning in PyTorch
Michela Paganini, Jessica Forde
In order to contrast the explosion in size of state-of-the-art machine learning models that can be attributed to the empirical advantages of over-parametrization, and due to the ne…