2 citations · 2 across the 2 of their papers we have counts for
6 papers
Position: Use Sparse Autoencoders to Discover Unknowns
Kenny Peng, Rajiv Movva, Jon Kleinberg +2
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptua…
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
Rajiv Movva, Smitha Milli, Sewon Min +1
Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback data encodes. While prior work studies…
Sparse Autoencoders for Hypothesis Generation
Rajiv Movva, Kenny Peng, Nikhil Garg +2
We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three…
Using large language models to promote health equity
Emma Pierson, Divya Shanmugam, Rajiv Movva +12
Advances in large language models (LLMs) have driven an explosion of interest about their societal impacts. Much of the discourse around how they will impact social equity has been…
Generative AI in Medicine
Divya Shanmugam, Monica Agrawal, Rajiv Movva +4
The increased capabilities of generative AI have dramatically expanded its possible use cases in medicine. We provide a comprehensive overview of generative AI use cases for clinic…
Annotation alignment: Comparing LLM and human annotations of conversational safety
Rajiv Movva, Pang Wei Koh, Emma Pierson
Do LLMs align with human perceptions of safety? We study this question via annotation alignment, the extent to which LLMs and humans agree when annotating the safety of user-chatbo…