7 papers
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
Elisabetta Rocchetti, Alfio Ferrara
Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-means (…
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
Elisabetta Rocchetti, Alfio Ferrara
Instruction tuning is commonly assumed to endow language models with a domain-general ability to follow instructions, yet the underlying mechanism remains poorly understood. Does i…
Unveiling Transformer Perception by Exploring Input Manifolds
Alessandro Benfenati, Alfio Ferrara, Alessio Marta +2
This paper introduces a general method for the exploration of equivalence classes in the input space of Transformer models. The proposed approach is based on sound mathematical the…
Modeling Transformers as complex networks to analyze learning dynamics
Elisabetta Rocchetti
The process by which Large Language Models (LLMs) acquire complex capabilities during training remains a key open question in mechanistic interpretability. This project investigate…
How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis
Elisabetta Rocchetti, Alfio Ferrara
Adhering to explicit length constraints, such as generating text with a precise word count, remains a significant challenge for Large Language Models (LLMs). This study aims at inv…
What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
Alfio Ferrara, Sergio Picascia, Laura Pinnavaia +3
Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on expli…