collaborators

9 papers

cs.LG2026

From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach

Nura Aljaafari, Danilo S. Carvalho, Andre Freitas

Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no s…

cs.CL2026

Emergence and Localisation of Semantic Role Circuits in LLMs

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Despite displaying semantic competence, large language models' internal mechanisms that ground abstract semantic structure remain insufficiently characterised. We propose a method…

cs.AI2025

elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings

Philip Osborne, Danilo S. Carvalho, André Freitas

We present elsciRL, an open-source Python library to facilitate the application of language solutions on reinforcement learning problems. We demonstrate the potential of our softwa…

cs.CL2025

TRACE: Training and Inference-Time Interpretability Analysis for Language Models

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Understanding when and how linguistic knowledge emerges during language model training remains a central challenge for interpretability. Most existing tools are post hoc, rely on s…

cs.CL2025

TRACE for Tracking the Emergence of Semantic Representations in Transformers

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Modern transformer models exhibit phase transitions during training, distinct shifts from memorisation to abstraction, but the mechanisms underlying these transitions remain poorly…

cs.CL2025

CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Large language models (LLMs) struggle with compositional generalisation, limiting their ability to systematically combine learned components to interpret novel inputs. While archit…