5 papers
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
Advey Nandan, Cheng-Ting Chou, Amrit Kurakula +4
We investigate the phenomenon of neuron universality in independently trained GPT-2 Small models, examining these universal neurons-neurons with consistently correlated activations…
Interpreting the Latent Structure of Operator Precedence in Language Models
Dharunish Yugeswardeenoo, Harshil Nukala, Ved Shah +4
Large Language Models (LLMs) have demonstrated impressive reasoning capabilities but continue to struggle with arithmetic tasks. Prior works largely focus on outputs or prompting s…
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
Cheng-Ting Chou, George Liu, Jessica Sun +4
Deterministically controlling the target generation language of large multilingual language models (LLMs) remains a fundamental challenge, particularly in zero-shot settings where…
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
Daniel Son, Sanjana Rathore, Andrew Rufail +6
We investigate feature universality in Gemma-2 language models (Gemma-2-2B and Gemma-2-9B), asking whether models with a four-fold difference in scale still converge on comparable…
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
Stanley Yu, Vaidehi Bulusu, Oscar Yasunaga +5
Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represen…