5 papers
Pathways of Visual Information Flow in Vision-Language Models
Israfel Salazar, Stella Frank, Dan Oneata +2
We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two dist…
Real-Time Progress Prediction in Reasoning Language Models
Hans Peter Lyngsøe Raaschou-Jensen, Constanza Fierro, Anders Søgaard
Recent reasoning language models, particularly those that employ long latent chains of thought, achieve strong performance on complex agentic tasks. However, as these models operat…
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…
Steering Language Models with Weight Arithmetic
Constanza Fierro, Fabien Roger
Providing high-quality feedback to Large Language Models (LLMs) on a diverse training distribution can be difficult and expensive, and providing feedback only on a narrow distribut…
How Do Multilingual Language Models Remember Facts?
Constanza Fierro, Negar Foroutan, Desmond Elliott +1
Large Language Models (LLMs) store and retrieve vast amounts of factual knowledge acquired during pre-training. Prior research has localized and identified mechanisms behind knowle…