2 citations · 2 across the 2 of their papers we have counts for
8 papers
Position: Use Sparse Autoencoders to Discover Unknowns
Kenny Peng, Rajiv Movva, Jon Kleinberg +2
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…
Mixing times of one-sided -transposition shuffles
Evita Nestoridi, Kenny Peng, Bryan Wong
We study mixing times of the one-sided -transposition shuffle. We prove that this shuffle mixes relatively slowly, even for big. Using the recent ``lifting eigenvectors'' te…
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
Nikhil Garg, Jon Kleinberg, Kenny Peng
We introduce a mathematical framework for the linear representation hypothesis (LRH), which asserts that intermediate layers of language models store features linearly. We separate…
Correlated Errors in Large Language Models
Elliot Kim, Avi Garg, Kenny Peng +1
Diversity in training data, architecture, and providers is assumed to mitigate homogeneity in LLMs. However, we lack empirical evidence on whether different LLMs differ meaningfull…
Sparse Autoencoders for Hypothesis Generation
Rajiv Movva, Kenny Peng, Nikhil Garg +2
We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…
A No Free Lunch Theorem for Human-AI Collaboration
Kenny Peng, Nikhil Garg, Jon Kleinberg
The gold standard in human-AI collaboration is complementarity -- when combined performance exceeds both the human and algorithm alone. We investigate this challenge in binary clas…