11 papers · 1 filter
Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
Nura Aljaafari, Andre Freitas
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introd…
Principles of Concept Representation in Sentence Encoders
Isabelle Mohr, John Dujany, Jonathan Souquet +1
What makes a sentence encoder produce good concept representations? We approach this through the lens of representational compositionality: an encoder supports a concept family onl…
Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder
Yingji Zhang, Danilo S. Carvalho, André Freitas
Integrating compositional and symbolic properties into current distributional semantic spaces can enhance the interpretability, controllability, compositionality, and generalisatio…
Emergence and Localisation of Semantic Role Circuits in LLMs
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Despite displaying semantic competence, large language models' internal mechanisms that ground abstract semantic structure remain insufficiently characterised. We propose a method…
Learning to Disentangle Latent Reasoning Rules with Language VAEs: A Systematic Study
Yingji Zhang, Marco Valentino, Danilo S. Carvalho +1
Incorporating explicit reasoning rules within the latent space of language models (LMs) offers a promising pathway to enhance generalisation, interpretability, and controllability.…
TRACE: Training and Inference-Time Interpretability Analysis for Language Models
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Understanding when and how linguistic knowledge emerges during language model training remains a central challenge for interpretability. Most existing tools are post hoc, rely on s…