activity
20172022
most citedScaling Language Models: Methods, Analysis & Insights from Training Gopher

243 citations · 402 across the 12 of their papers we have counts for

collaborators

30 papers

cs.CL202223 cited

Unified Scaling Laws for Routed Language Models

Aidan Clark, Diego de las Casas, Aurelia Guy +23

The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: archite…

cs.CL2022243 cited

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…

cs.CV2020

Prune Responsibly

Michela Paganini

Irrespective of the specific definition of fairness in a machine learning application, pruning the underlying model affects it. We investigate and document the emergence and exacer…

cs.LG20202 cited

Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding

Michela Paganini, Jessica Zosa Forde

The observation of sparse trainable sub-networks within over-parametrized networks - also known as Lottery Tickets (LTs) - has prompted inquiries around their trainability, scaling…

cs.SE20201 cited

dagger: A Python Framework for Reproducible Machine Learning Experiment Orchestration

Michela Paganini, Jessica Zosa Forde

Many research directions in machine learning, particularly in deep learning, involve complex, multi-stage experiments, commonly involving state-mutating operations acting on models…

cs.LG20207 cited

Streamlining Tensor and Network Pruning in PyTorch

Michela Paganini, Jessica Forde

In order to contrast the explosion in size of state-of-the-art machine learning models that can be attributed to the empirical advantages of over-parametrization, and due to the ne…