8 papers
Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
Nura Aljaafari, Andre Freitas
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introd…
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
Nura Aljaafari, Marco Valentino, André Freitas
Predicting a label correctly does not necessarily require representing the operation that produces it. Transformer representations are known to carry label-level information, but w…
From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach
Nura Aljaafari, Danilo S. Carvalho, Andre Freitas
Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no s…
Emergence and Localisation of Semantic Role Circuits in LLMs
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Despite displaying semantic competence, large language models' internal mechanisms that ground abstract semantic structure remain insufficiently characterised. We propose a method…
TRACE: Training and Inference-Time Interpretability Analysis for Language Models
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Understanding when and how linguistic knowledge emerges during language model training remains a central challenge for interpretability. Most existing tools are post hoc, rely on s…
TRACE for Tracking the Emergence of Semantic Representations in Transformers
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Modern transformer models exhibit phase transitions during training, distinct shifts from memorisation to abstraction, but the mechanisms underlying these transitions remain poorly…