2 citations · 2 across the 2 of their papers we have counts for
7 papers
Position: Use Sparse Autoencoders to Discover Unknowns
Kenny Peng, Rajiv Movva, Jon Kleinberg +2
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…
Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift
Kevin Ren, Manish Raghavan, Nikhil Garg
Deployed approaches for AI text detection often rely on training-time access to labeled datasets of both human-written and AI-generated text. This approach is vulnerable to three t…
The Subjectivity of Monoculture
Nathanael Jo, Nikhil Garg, Manish Raghavan
Machine learning models -- including large language models (LLMs) -- are often said to exhibit monoculture, where outputs agree strikingly often. But what does it actually mean for…
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
Nikhil Garg, Jon Kleinberg, Kenny Peng
We introduce a mathematical framework for the linear representation hypothesis (LRH), which asserts that intermediate layers of language models store features linearly. We separate…
Correlated Errors in Large Language Models
Elliot Kim, Avi Garg, Kenny Peng +1
Diversity in training data, architecture, and providers is assumed to mitigate homogeneity in LLMs. However, we lack empirical evidence on whether different LLMs differ meaningfull…
Sparse Autoencoders for Hypothesis Generation
Rajiv Movva, Kenny Peng, Nikhil Garg +2
We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…