4 papers
From superposition to sparse codes: interpretable representations in neural networks
David Klindt, Charles O'Neill, Patrik Reizinger +2
Understanding how information is represented in neural networks is a fundamental challenge in both neuroscience and artificial intelligence. Despite their nonlinear architectures,…
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
Charles O'Neill, Alim Gumran, David Klindt
A recent line of work has shown promise in using sparse autoencoders (SAEs) to uncover interpretable features in neural network representations. However, the simple linear-nonlinea…
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
Charles O'Neill
Self-attention mechanisms have revolutionised deep learning architectures, yet their core mathematical structures remain incompletely understood. In this work, we develop a categor…
Disentangling Dense Embeddings with Sparse Autoencoders
Charles O'Neill, Christine Ye, Kartheik Iyer +1
Sparse autoencoders (SAEs) have shown promise in extracting interpretable features from complex neural networks. We present one of the first applications of SAEs to dense text embe…