4 papers
The Value Axis: Language Models Encode Whether They're on the Right Track
Nick Jiang, Isaac Kauvar, Jack Lindsey
We investigate whether language models internally track the value of their current trajectory, defined as the likelihood that their ongoing strategy will achieve their goals. Using…
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Nick Jiang, Xiaoqing Sun, Lisa Dunlap +2
Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data. Current metho…
Vision Transformers Don't Need Trained Registers
Nick Jiang, Amil Dravid, Alexei Efros +1
We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers - the emergence of high-norm tokens that lead to noisy attention maps (Darcet et a…
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Nick Jiang, Anish Kachinthaya, Suzie Petryk +1
We investigate the internal representations of vision-language models (VLMs) to address hallucinations, a persistent challenge despite advances in model size and training. We proje…