collaborators

7 papers

cs.AI2026

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

Elisabetta Rocchetti, Alfio Ferrara

Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-means (…

cs.AI2026

How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism

Elisabetta Rocchetti, Alfio Ferrara

Instruction tuning is commonly assumed to endow language models with a domain-general ability to follow instructions, yet the underlying mechanism remains poorly understood. Does i…

cs.LG2025

Unveiling Transformer Perception by Exploring Input Manifolds

Alessandro Benfenati, Alfio Ferrara, Alessio Marta +2

This paper introduces a general method for the exploration of equivalence classes in the input space of Transformer models. The proposed approach is based on sound mathematical the…

cs.LG2025

Modeling Transformers as complex networks to analyze learning dynamics

Elisabetta Rocchetti

The process by which Large Language Models (LLMs) acquire complex capabilities during training remains a key open question in mechanistic interpretability. This project investigate…

cs.CL2025

How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis

Elisabetta Rocchetti, Alfio Ferrara

Adhering to explicit length constraints, such as generating text with a precise word count, remains a significant challenge for Large Language Models (LLMs). This study aims at inv…

cs.CL2025

What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content

Alfio Ferrara, Sergio Picascia, Laura Pinnavaia +3

Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on expli…