12 citations · 14 across the 10 of their papers we have counts for
11 papers · 1 filter
Latent-Space Intervention for Cross-Lingual Factual Consistency: Consistency Improvements without Accuracy Drops
Faeze Ghorbanpour, Constanza Fierro, Alexander Fraser +1
Large Language Models (LLMs) often answer the same factual question differently across languages. We study whether cross-lingual latent-space intervention can reduce this inconsist…
Steering Language Models with Weight Arithmetic
Constanza Fierro, Fabien Roger
Providing high-quality feedback to Large Language Models (LLMs) on a diverse training distribution can be difficult and expensive, and providing feedback only on a narrow distribut…
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…
Defining Knowledge: Bridging Epistemology and Large Language Models
Constanza Fierro, Ruchira Dhar, Filippos Stamatiou +2
Knowledge claims are abundant in the literature on large language models (LLMs); but can we say that GPT-4 truly "knows" the Earth is round? To address this question, we review sta…
How Do Multilingual Language Models Remember Facts?
Constanza Fierro, Negar Foroutan, Desmond Elliott +1
Large Language Models (LLMs) store and retrieve vast amounts of factual knowledge acquired during pre-training. Prior research has localized and identified mechanisms behind knowle…
MuLan: A Study of Fact Mutability in Language Models
Constanza Fierro, Nicolas Garneau, Emanuele Bugliarello +2
Facts are subject to contingencies and can be true or false in different circumstances. One such contingency is time, wherein some facts mutate over a given period, e.g., the presi…