collaborators

5 papers

cs.LG2025

Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact

Advey Nandan, Cheng-Ting Chou, Amrit Kurakula +4

We investigate the phenomenon of neuron universality in independently trained GPT-2 Small models, examining these universal neurons-neurons with consistently correlated activations…

cs.CL2025

Interpreting the Latent Structure of Operator Precedence in Language Models

Dharunish Yugeswardeenoo, Harshil Nukala, Ved Shah +4

Large Language Models (LLMs) have demonstrated impressive reasoning capabilities but continue to struggle with arithmetic tasks. Prior works largely focus on outputs or prompting s…

cs.CL2025

Causal Language Control in Multilingual Transformers via Sparse Feature Steering

Cheng-Ting Chou, George Liu, Jessica Sun +4

Deterministically controlling the target generation language of large multilingual language models (LLMs) remains a fundamental challenge, particularly in zero-shot settings where…

cs.CL2025

Semantic Convergence: Investigating Shared Representations Across Scaled LLMs

Daniel Son, Sanjana Rathore, Andrew Rufail +6

We investigate feature universality in Gemma-2 language models (Gemma-2-2B and Gemma-2-9B), asking whether models with a four-fold difference in scale still converge on comparable…

cs.LG2025

From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs

Stanley Yu, Vaidehi Bulusu, Oscar Yasunaga +5

Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represen…