15 papers · 1 filter
A Mechanistic Understanding of Pronoun Fidelity in LLMs
Katharina Trinley, Jesujoba O. Alabi, Dietrich Klakow +1
Faithful and robust pronoun use is important for fair and coherent generations, yet large language models largely fail when multiple referents use different pronouns. To study the…
Your Multimodal Speech Model Says I Have a Face for Radio
Maya K. Nachesa, Vlad Niculae, Vagrant Gautam
As large neural models have become better at language tasks, researchers are increasingly building multi- and omnimodal models that handle more modalities of data. One example is t…
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
Fabian Mewes, Anne Lauscher, Vagrant Gautam
Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about reference. More recently, the interp…
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
Jakob Schuster, Vagrant Gautam, Katja Markert
As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their behavior under knowledge conflicts.…
Teaching and Critiquing Conceptualization and Operationalization in NLP
Vagrant Gautam
NLP researchers regularly invoke abstract concepts like "interpretability," "bias," "reasoning," and "stereotypes," without defining them. Each subfield has a shared understanding…
Aligned Probing: Relating Toxic Behavior and Model Internals
Andreas Waldis, Vagrant Gautam, Anne Lauscher +2
We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (inte…