32 citations · 45 across the 9 of their papers we have counts for
10 papers · 1 filter
Orthogonal Self-Attention
Leo Zhang, James Martens
Softmax Self-Attention (SSA) is a key component of Transformer architectures. However, when utilised within skipless architectures, which aim to improve representation learning, re…
Cutting the Skip: Training Residual-Free Transformers
Yiping Ji, James Martens, Jianqiao Zheng +5
Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without skip (residual) connectio…
Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation
Bobby He, James Martens, Guodong Zhang +4
Skip connections and normalisation layers form two standard architectural components that are ubiquitous for the training of Deep Neural Networks (DNNs), but whose precise roles ar…
Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers
Guodong Zhang, Aleksandar Botev, James Martens
Training very deep neural networks is still an extremely challenging task. The common solution is to use shortcut connections and normalization layers, which are both crucial ingre…
Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
James Martens, Andy Ballard, Guillaume Desjardins +4
Using an extended and formalized version of the Q/C map analysis of Poole et al. (2016), along with Neural Tangent Kernel theory, we identify the main pathologies present in deep n…
On the validity of kernel approximations for orthogonally-initialized neural networks
James Martens
In this note we extend kernel function approximation results for neural networks with Gaussian-distributed weights to single-layer networks initialized using Haar-distributed rando…