8 papers
Language Models Embody and Amplify Human Cognitive Distortions: What Is to Be Done?
Arnau Marin-Llobet, Steven A. Lehr, Mahzarin R. Banaji
Human judgment is fundamentally prone to error. A promise of AI is that it will rid decisions of bias and ensure a fairer and safer world for all. Yet research unequivocally demons…
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
Arnau Marin-Llobet, Stefan Heimersheim
A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant circuit-finding approaches focus on a spe…
Can neurons speak? Semantic narration of vision at single-cell resolution
Arnau Marin-Llobet, Richard Hakim, Sara Matias +3
Identifying what individual neurons encode in higher-order visual cortex is an open problem. Responses resist intuitive parameterization, and the deep-network embeddings used in th…
Vision-Language Models Suppress Female Representations Under Ambiguous Input
Arnau Marin-Llobet, Simon Henniger, Mahzarin R. Banaji
Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous i…
LiFT: Lifted Inter-slice Feature Trajectories for 3D Image Generation from 2D Generators
Xinhe Zhang, Yuyang Zhang, Pengfei Jin +3
High-resolution 3D medical image generation remains challenging because fully volumetric models are computationally expensive, while efficient 2D slice generators often fail to pre…
Automated Interpretability and Feature Discovery in Language Models with Agents
Arnau Marin-Llobet, Javier Ferrando
We introduce an autonomous multiagent framework for mechanistic interpretability that automates both explaining and finding internal features in large language models. The system r…