Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
Elisabetta Rocchetti, Alfio Ferrara
Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-means (…
cs.AI2026
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
Elisabetta Rocchetti, Alfio Ferrara
Instruction tuning is commonly assumed to endow language models with a domain-general ability to follow instructions, yet the underlying mechanism remains poorly understood. Does i…