3 papers
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…
cs.LG2024
How Reliable are Causal Probing Interventions?
Marc Canby, Adam Davies, Chirag Rastogi +1
Causal probing aims to analyze foundation models by examining how intervening on their representation of various latent properties impacts their outputs. Recent works have cast dou…
cs.CL2023
A Framework for Bidirectional Decoding: Case Study in Morphological Inflection
Marc E. Canby, Julia Hockenmaier
Transformer-based encoder-decoder models that generate outputs in a left-to-right fashion have become standard for sequence-to-sequence tasks. In this paper, we propose a framework…