58 citations · 116 across the 36 of their papers we have counts for
13 papers · 1 filter
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
Christopher Warner, Jonas Mago, JR Huml +1
We introduce ZUNA1.1, a 380M-parameter diffusion autoencoder for flexible EEG signal reconstruction. ZUNA1.1 is capable of reconstructing variable length sequences of up to 30s, wi…
Scaling Adaptive Depth with Norm-Agnostic Residual Networks
Tomás Figliolia, Beren Millidge
Residual architectures are ubiquitous in deep learning, but they suffer from a subtle structural limitation: the norm of the residual stream can grow rapidly with depth. As a resul…
Hybrid Associative Memories
Leon Lufkin, Tomás Figliolia, Beren Millidge +1
Recurrent neural networks (RNNs) and self-attention are both widely used sequence-mixing layers that maintain an internal memory. However, this memory is constructed using two orth…
Generalising E-prop to Deep Networks
Beren Millidge
Recurrent networks are typically trained with backpropagation through time (BPTT). However, BPTT requires storing the history of all states in the network and then replaying them s…
The Zamba2 Suite: Technical Report
Paolo Glorioso, Quentin Anthony, Yury Tokpanov +5
In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance again…
Interpreting Neural Networks through the Polytope Lens
Sid Black, Lee Sharkey, Leo Grinsztajn +8
Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Pre…