From the 1 of 6 linked papers with an AI index.
6 papers
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Robert Graham, Edward Stevinson, Yariv Barsheshat
The paper shows that fine‑tuning large language models on small, factually defensible datasets can cause broad ideological shifts across unrelated topics, and introduces metrics to…
Adversarial Attacks Leverage Interference Between Features in Superposition
Edward Stevinson, Lucas Prieto, Melih Barsbey +1
Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decisio…
From Data Statistics to Feature Geometry: How Correlations Shape Superposition
Lucas Prieto, Edward Stevinson, Melih Barsbey +2
A central idea in mechanistic interpretability is that neural networks represent more features than they have dimensions, arranging them in superposition to form an over-complete b…
ContextBench: Modifying Contexts for Targeted Latent Activation
Robert Graham, Edward Stevinson, Leo Richter +3
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…
A Scalable Approach to Probabilistic Neuro-Symbolic Robustness Verification
Vasileios Manginas, Nikolaos Manginas, Edward Stevinson +4
Neuro-Symbolic Artificial Intelligence (NeSy AI) has emerged as a promising direction for integrating neural learning with symbolic reasoning. Typically, in the probabilistic varia…
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
Sonia Joseph, Praneet Suresh, Lorenz Hufe +7
Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…