Publications (9)
Towards Spectroscopy: Susceptibility Clusters in Language Models
Andrew Gordon, Garrett Baker, George Wang +3
Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distrib…
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
Sean P. Fillingham, Andrew Gordon, Peter Lai +3
Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs…
Rows from Many Sources: Enriching row completions from Wikidata with a pre-trained Language Model
Carina Negreanu, Alperen Karaoglu, Jack Williams +4
Row completion is the task of augmenting a given table of text and numbers with additional, relevant rows. The task divides into two steps: subject suggestion, the task of populati…
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
Nora Petrova, Andrew Gordon, Enzo Blindow
The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from…
Recent Spin Results from STAR
Andrew Gordon
In Run 8 at RHIC, STAR significantly enhanced its forward acceptance relative to previous years with the commissioning of a new detector, the Forward Meson Spectrometer (FMS). The…
The Wreath Process: A totally generative model of geometric shape based on nested symmetries
Diana Borsa, Thore Graepel, Andrew Gordon
We consider the problem of modelling noisy but highly symmetric shapes that can be viewed as hierarchies of whole-part relationships in which higher level objects are composed of t…
Embryology of a Language Model
George Wang, Garrett Baker, Andrew Gordon +1
Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistica…
Current Status of Transverse Spin at STAR
Andrew Gordon
The origins of the proton spin remain an area of active investigation. The Relativistic Heavy Ion Collider (RHIC) at Brookhaven National Lab uniquely provides polarized proton data…
Prevalence and prevention of large language model use in crowd work
Veniamin Veselovsky, Manoel Horta Ribeiro, Philip Cozzolino +3
We show that the use of large language models (LLMs) is prevalent among crowd workers, and that targeted mitigation strategies can significantly reduce, but not eliminate, LLM use.…