works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.LG2026

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

Robert Graham, Edward Stevinson, Yariv Barsheshat

The paper shows that fine‑tuning large language models on small, factually defensible datasets can cause broad ideological shifts across unrelated topics, and introduces metrics to…

cs.LG2026

Adversarial Attacks Leverage Interference Between Features in Superposition

Edward Stevinson, Lucas Prieto, Melih Barsbey +1

Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decisio…

cs.LG2026

From Data Statistics to Feature Geometry: How Correlations Shape Superposition

Lucas Prieto, Edward Stevinson, Melih Barsbey +2

A central idea in mechanistic interpretability is that neural networks represent more features than they have dimensions, arranging them in superposition to form an over-complete b…

cs.AI2026

ContextBench: Modifying Contexts for Targeted Latent Activation

Robert Graham, Edward Stevinson, Leo Richter +3

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…

cs.AI2025

A Scalable Approach to Probabilistic Neuro-Symbolic Robustness Verification

Vasileios Manginas, Nikolaos Manginas, Edward Stevinson +4

Neuro-Symbolic Artificial Intelligence (NeSy AI) has emerged as a promising direction for integrating neural learning with symbolic reasoning. Typically, in the probabilistic varia…

cs.CV2025

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Sonia Joseph, Praneet Suresh, Lorenz Hufe +7

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…