3 papers
cs.LG2026
Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability
David N. Olivieri, Antonio F. Pérez RodrÃguez
Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering dire…
cs.AI2026
Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents
David N. Olivieri, Roque J. Hernández
Scientific theory shift in AI agents requires more than fitting equations to data. An artificial scientific agent must detect whether an existing representational framework remains…
cs.AI2026
Patch-Effect Graph Kernels for LLM Interpretability
Ruben Fernandez-Boullon, David N. Olivieri
Mechanistic interpretability aims to reverse-engineer transformer computations by identifying causal circuits through activation patching. However, scaling these interventions acro…