3 citations · 8 across the 6 of their papers we have counts for
5 papers · 1 filter
Optimal Formats for Weight Quantisation
Douglas Orr, Luka Ribar, Carlo Luschi
Weight quantisation is an essential technique for enabling efficient training and deployment of modern deep learning models. However, the recipe book of quantisation formats is lar…
u-P: The Unit-Scaled Maximal Update Parametrization
Charlie Blake, Constantin Eichenberg, Josef Dean +7
The Maximal Update Parametrization (P) aims to make the optimal hyperparameters (HPs) of a model independent of its size, allowing them to be swept using a cheap proxy model rat…
SparQ Attention: Bandwidth-Efficient LLM Inference
Luka Ribar, Ivan Chelombiev, Luke Hudlass-Galley +3
The computational difficulties of large language model (LLM) inference remain a significant obstacle to their widespread deployment. The need for many applications to support long…
PopSparse: Accelerated block sparse matrix multiplication on IPU
Zhiyi Li, Douglas Orr, Valeriu Ohan +5
Reducing the computational cost of running large scale neural networks using sparsity has attracted great attention in the deep learning community. While much success has been achi…
BESS: Balanced Entity Sampling and Sharing for Large-Scale Knowledge Graph Completion
Alberto Cattaneo, Daniel Justus, Harry Mellor +7
We present the award-winning submission to the WikiKG90Mv2 track of OGB-LSC@NeurIPS 2022. The task is link-prediction on the large-scale knowledge graph WikiKG90Mv2, consisting of…