2 citations · 2 across the 1 of their papers we have counts for
4 papers
Position: Use Sparse Autoencoders to Discover Unknowns
Kenny Peng, Rajiv Movva, Jon Kleinberg +2
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
Nikhil Garg, Jon Kleinberg, Kenny Peng
We introduce a mathematical framework for the linear representation hypothesis (LRH), which asserts that intermediate layers of language models store features linearly. We separate…
Sparse Autoencoders for Hypothesis Generation
Rajiv Movva, Kenny Peng, Nikhil Garg +2
We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…
A No Free Lunch Theorem for Human-AI Collaboration
Kenny Peng, Nikhil Garg, Jon Kleinberg
The gold standard in human-AI collaboration is complementarity -- when combined performance exceeds both the human and algorithm alone. We investigate this challenge in binary clas…