activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Emergence and Localisation of Semantic Role Circuits in LLMs

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Despite displaying semantic competence, large language models' internal mechanisms that ground abstract semantic structure remain insufficiently characterised. We propose a method…

cs.CL2025

TRACE: Training and Inference-Time Interpretability Analysis for Language Models

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Understanding when and how linguistic knowledge emerges during language model training remains a central challenge for interpretability. Most existing tools are post hoc, rely on s…

cs.CL2025

TRACE for Tracking the Emergence of Semantic Representations in Transformers

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Modern transformer models exhibit phase transitions during training, distinct shifts from memorisation to abstraction, but the mechanisms underlying these transitions remain poorly…

cs.CL2025

CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Large language models (LLMs) struggle with compositional generalisation, limiting their ability to systematically combine learned components to interpret novel inputs. While archit…

cs.CL2025

Interpreting token compositionality in LLMs: A robustness analysis

Nura Aljaafari, Danilo S. Carvalho, André Freitas

Understanding the internal mechanisms of large language models (LLMs) is integral to enhancing their reliability, interpretability, and inference processes. We present Constituent-…

cs.CL2025

Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis

João Pedro Gandarela, Danilo S. Carvalho, André Freitas

This work presents a novel systematic methodology to analyse the capabilities and limitations of Large Language Models (LLMs) with feedback from a formal inference engine, on logic…