44 citations · 198 across the 13 of their papers we have counts for
24 papers
Relative Positional Encoding for Transformers with Linear Complexity
Antoine Liutkus, Ondřej Cífka, Shih-Lun Wu +3
Recent advances in Transformer models allow for unprecedented sequence lengths, due to linear space and time complexity. In the meantime, relative positional encoding (RPE) was pro…
Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks
Melih Barsbey, Milad Sefidgaran, Murat A. Erdogdu +2
Neural network compression techniques have become increasingly popular as they can drastically reduce the storage and computation requirements for very large networks. Recent empir…
Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
Alexander Camuto, Xiaoyu Wang, Lingjiong Zhu +3
Gaussian noise injections (GNIs) are a family of simple and widely-used regularisation methods for training neural networks, where one injects additive or multiplicative Gaussian n…
Self-Supervised VQ-VAE for One-Shot Music Style Transfer
Ondřej Cífka, Alexey Ozerov, Umut Şimşekli +1
Neural style transfer, allowing to apply the artistic style of one image to another, has become one of the most widely showcased computer vision applications shortly after its intr…
Quantitative Propagation of Chaos for SGD in Wide Neural Networks
Valentin De Bortoli, Alain Durmus, Xavier Fontaine +1
In this paper, we investigate the limiting behavior of a continuous-time counterpart of the Stochastic Gradient Descent (SGD) algorithm applied to two-layer overparameterized neura…
Explicit Regularisation in Gaussian Noise Injections
Alexander Camuto, Matthew Willetts, Umut Şimşekli +2
We study the regularisation induced in neural networks by Gaussian noise injections (GNIs). Though such injections have been extensively studied when applied to data, there have be…