5 papers
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
Mikkel Godsk Jørgensen, Lars Kai Hansen
Sparse Autoencoders (SAEs) have been seen as a promising avenue for exploring the internals of Large Language Models (LLMs) and for steering model output generation. When AxBench -…
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa +10
EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply To…
Pretraining on Sleep Data Improves non-Sleep Biosignal Tasks
William Lehn-Schiøler, Magnus Ruud Kjær, Phillip Hempel +7
Sleep foundation models have recently demonstrated strong performance on in-domain polysomnography tasks, including sleep staging, apnea detection, and disease risk prediction. In…
From Colors to Classes: Emergence of Concepts in Vision Transformers
Teresa Dorszewski, Lenka TÄtková, Robert Jenssen +2
Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs p…
How Redundant Is the Transformer Stack in Speech Representation Models?
Teresa Dorszewski, Albert Kjøller Jacobsen, Lenka TÄtková +1
Self-supervised speech representation models, particularly those leveraging transformer architectures, have demonstrated remarkable performance across various tasks such as speech…