activity
20242026
most citedPosition: Use Sparse Autoencoders to Discover Unknowns

2 citations · 2 across the 2 of their papers we have counts for

collaborators

6 papers

cs.LG20262 cited

Position: Use Sparse Autoencoders to Discover Unknowns

Kenny Peng, Rajiv Movva, Jon Kleinberg +2

While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…

cs.CL2026

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

Rajiv Movva, Smitha Milli, Sewon Min +1

Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback data encodes. While prior work studies…

cs.CL2025

Sparse Autoencoders for Hypothesis Generation

Rajiv Movva, Kenny Peng, Nikhil Garg +2

We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…

cs.CY2025

Using large language models to promote health equity

Emma Pierson, Divya Shanmugam, Rajiv Movva +12

Advances in large language models (LLMs) have driven an explosion of interest about their societal impacts. Much of the discourse around how they will impact social equity has been…

cs.LG2024

Generative AI in Medicine

Divya Shanmugam, Monica Agrawal, Rajiv Movva +4

The increased capabilities of generative AI have dramatically expanded its possible use cases in medicine. We provide a comprehensive overview of generative AI use cases for clinic…

cs.CL2024

Annotation alignment: Comparing LLM and human annotations of conversational safety

Rajiv Movva, Pang Wei Koh, Emma Pierson

Do LLMs align with human perceptions of safety? We study this question via annotation alignment, the extent to which LLMs and humans agree when annotating the safety of user-chatbo…