9 papers
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
Nurkhan Laiyk, Gerard I. Gállego, Javier Ferrando +1
Function vectors (FVs) are vector representations of tasks extracted from model activations during in-context learning. While prior work has shown that multilingual model represent…
Automated Interpretability and Feature Discovery in Language Models with Agents
Arnau Marin-Llobet, Javier Ferrando
We introduce an autonomous multiagent framework for mechanistic interpretability that automates both explaining and finding internal features in large language models. The system r…
Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
Sergi Masip, Gido M. van de Ven, Javier Ferrando +1
Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying mechanisms. We introduce a mechani…
Weight space Detection of Backdoors in LoRA Adapters
David Puertolas Merenciano, Ekaterina Vasyagina, Kevin Zhu +2
LoRA adapters let users fine-tune large language models (LLMs) efficiently. However, LoRA adapters are shared through open repositories like Hugging Face Hub \citep{huggingface_hub…
Language Models Can Explain Visual Features via Steering
Javier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres +3
Sparse Autoencoders uncover thousands of features in vision models, yet explaining these features without requiring human intervention remains an open challenge. While previous wor…
Real-Time Detection of Hallucinated Entities in Long-Form Generation
Oscar Obeso, Andy Arditi, Javier Ferrando +3
Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing halluc…