8 papers
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
Elisabetta Rocchetti, Alfio Ferrara
Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-means (…
Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models
Luigi Curini, Alfio Ferrara, Giovanni Pagano +1
Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts t…
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
Elisabetta Rocchetti, Alfio Ferrara
Instruction tuning is commonly assumed to endow language models with a domain-general ability to follow instructions, yet the underlying mechanism remains poorly understood. Does i…
Quid est VERITAS? A Modular Framework for Archival Document Analysis
Leonardo Bassanini, Ludovico Biancardi, Alfio Ferrara +3
The digitisation of historical documents has traditionally been conceived as a process limited to character-level transcription, producing flat text that lacks the structural and s…
Unveiling Transformer Perception by Exploring Input Manifolds
Alessandro Benfenati, Alfio Ferrara, Alessio Marta +2
This paper introduces a general method for the exploration of equivalence classes in the input space of Transformer models. The proposed approach is based on sound mathematical the…
How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis
Elisabetta Rocchetti, Alfio Ferrara
Adhering to explicit length constraints, such as generating text with a precise word count, remains a significant challenge for Large Language Models (LLMs). This study aims at inv…