activity
20242026
most citedPosition: Use Sparse Autoencoders to Discover Unknowns

2 citations · 2 across the 2 of their papers we have counts for

collaborators

8 papers

cs.LG20262 cited

Position: Use Sparse Autoencoders to Discover Unknowns

Kenny Peng, Rajiv Movva, Jon Kleinberg +2

While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…

math.PR2026

Mixing times of one-sided -transposition shuffles

Evita Nestoridi, Kenny Peng, Bryan Wong

We study mixing times of the one-sided -transposition shuffle. We prove that this shuffle mixes relatively slowly, even for big. Using the recent ``lifting eigenvectors'' te…

cs.LG2026

How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?

Nikhil Garg, Jon Kleinberg, Kenny Peng

We introduce a mathematical framework for the linear representation hypothesis (LRH), which asserts that intermediate layers of language models store features linearly. We separate…

cs.CL2025

Correlated Errors in Large Language Models

Elliot Kim, Avi Garg, Kenny Peng +1

Diversity in training data, architecture, and providers is assumed to mitigate homogeneity in LLMs. However, we lack empirical evidence on whether different LLMs differ meaningfull…

cs.CL2025

Sparse Autoencoders for Hypothesis Generation

Rajiv Movva, Kenny Peng, Nikhil Garg +2

We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…

cs.AI2024

A No Free Lunch Theorem for Human-AI Collaboration

Kenny Peng, Nikhil Garg, Jon Kleinberg

The gold standard in human-AI collaboration is complementarity -- when combined performance exceeds both the human and algorithm alone. We investigate this challenge in binary clas…