5 papers
Interpreting the Latent Structure of Operator Precedence in Language Models
Dharunish Yugeswardeenoo, Harshil Nukala, Ved Shah +4
Large Language Models (LLMs) have demonstrated impressive reasoning capabilities but continue to struggle with arithmetic tasks. Prior works largely focus on outputs or prompting s…
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
Advey Nandan, Cheng-Ting Chou, Amrit Kurakula +4
We investigate the phenomenon of neuron universality in independently trained GPT-2 Small models, examining these universal neurons-neurons with consistently correlated activations…
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
Daniel Son, Sanjana Rathore, Andrew Rufail +6
We investigate feature universality in Gemma-2 language models (Gemma-2-2B and Gemma-2-9B), asking whether models with a four-fold difference in scale still converge on comparable…
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
Cheng-Ting Chou, George Liu, Jessica Sun +4
Deterministically controlling the target generation language of large multilingual language models (LLMs) remains a fundamental challenge, particularly in zero-shot settings where…
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
Stanley Yu, Vaidehi Bulusu, Oscar Yasunaga +5
Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represen…