activity
20212024
most citedLocalizing Model Behavior with Path Patching

4 citations · 7 across the 9 of their papers we have counts for

collaborators

9 papers

cs.LG20241 cited

pyvene: A Library for Understanding and Improving PyTorch Models via Interventions

Zhengxuan Wu, Atticus Geiger, Aryaman Arora +5

Interventions on model-internal states are fundamental operations in many areas of AI, including model editing, steering, robustness, and interpretability. To facilitate such resea…

cs.CL2024

CausalGym: Benchmarking causal interpretability methods on linguistic tasks

Aryaman Arora, Dan Jurafsky, Christopher Potts

Language models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons).…

eess.AS20242 cited

Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens

Nay San, Georgios Paraskevopoulos, Aryaman Arora +4

While massively multilingual speech models like wav2vec 2.0 XLSR-128 can be directly fine-tuned for automatic speech recognition (ASR), downstream performance can still be relative…

cs.LG2024

A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments

Zhengxuan Wu, Atticus Geiger, Jing Huang +4

We respond to the recent paper by Makelov et al. (2023), which reviews subspace interchange intervention methods like distributed alignment search (DAS; Geiger et al. 2023) and cla…

cs.CL2023

IruMozhi: Automatically classifying diglossia in Tamil

Kabilan Prasanna, Aryaman Arora

Tamil, a Dravidian language of South Asia, is a highly diglossic language with two very different registers in everyday use: Literary Tamil (preferred in writing and formal communi…

cs.CL2023

Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP

Vedant Palit, Rohan Pandey, Aryaman Arora +1

Mechanistic interpretability seeks to understand the neural mechanisms that enable specific behaviors in Large Language Models (LLMs) by leveraging causality-based methods. While t…