Publications (14)
Nonorientable four-ball genus can be arbitrarily large
Joshua Batson
The nonorientable four-ball genus of a knot K is the smallest first Betti number of any smoothly embedded, nonorientable surface F in B^4 bounding K. In contrast to the orientable…
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…
Patch2Self: Denoising Diffusion MRI with Self-Supervised Learning
Shreyas Fadnavis, Joshua Batson, Eleftherios Garyfallidis
Diffusion-weighted magnetic resonance imaging (DWI) is the only noninvasive method for quantifying microstructure and reconstructing white-matter pathways in the living human brain…
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken +32
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…
A Link Splitting Spectral Sequence in Khovanov Homology
Joshua Batson, Cotton Seed
We construct a new spectral sequence beginning at the Khovanov homology of a link and converging to the Khovanov homology of the disjoint union of its components. The page at which…
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible re…
Topological Obstructions to Autoencoding
Joshua Batson, C. Grace Haaf, Yonatan Kahn +1
Autoencoders have been proposed as a powerful tool for model-independent anomaly detection in high-energy physics. The operating principle is that events which do not belong to the…
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus +23
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictiona…
Emotion Concepts and their Function in a Large Language Model
Nicholas Sofroniew, Isaac Kauvar, William Saunders +13
Large language models (LLMs) sometimes appear to exhibit emotional reactions. We investigate why this is the case in Claude Sonnet 4.5 and explore implications for alignment-releva…
Twice-Ramanujan Sparsifiers
Joshua Batson, Daniel A. Spielman, Nikhil Srivastava
We prove that every graph has a spectral sparsifier with a number of edges linear in its number of vertices. As linear-sized spectral sparsifiers of complete graphs are expanders,…
Scaling Laws in Jet Classification
Joshua Batson, Yonatan Kahn
We demonstrate the emergence of scaling laws in the benchmark top versus QCD jet classification problem in collider physics. Six distinct physically-motivated classifiers exhibit p…
Image Deconvolution via Noise-Tolerant Self-Supervised Inversion
Hirofumi Kobayashi, Ahmet Can Solak, Joshua Batson +1
We propose a general framework for solving inverse problems in the presence of noise that requires no signal prior, no noise estimate, and no clean training data. We only require t…
Noise2Self: Blind Denoising by Self-Supervision
Joshua Batson, Loic Royer
We propose a general framework for denoising high-dimensional measurements which requires no prior on the signal, no estimate of the noise, and no clean training data. The only ass…
When Models Manipulate Manifolds: The Geometry of a Counting Task
Wes Gurnee, Emmanuel Ameisen, Isaac Kauvar +4
Language models can perceive visual properties of text despite receiving only sequences of tokens-we mechanistically investigate how Claude 3.5 Haiku accomplishes one such task: li…