1 citations · 1 across the 7 of their papers we have counts for
8 papers
Entangled Representations Amplify Collateral Damage in Unlearning
Evžen Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt
A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning h…
Spatial Generalization Tests for Machine Learning-based Weather Models to Assess Physical Consistency
Maren Höver, Milan Klöwer, Christian Schroeder de Witt +1
Machine learning-based weather prediction is revolutionizing weather forecasting by learning from weather data in present-day climate. However, generalization to other climates rem…
A Note on the Strategic Confinement Problem
Christian Schroeder de Witt
Lampson's confinement problem asks how to prevent a program that processes confidential information from leaking it to a third party. We introduce the strategic confinement problem…
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
Aaron Rose, Carissa Cullen, Sahar Abdelnabi +3
As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on…
Chronos: The AI Co-Historian
Lorenz Hufe, Niclas Griesshaber, Gavin Greif +5
AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical research remains limited due to the lack of…
OpenSanctions Pairs: Large-Scale Entity Matching with LLMs
Chandler Smith, Magnus Sesodia, Friedrich Lindenberg +1
We release OpenSanctions Pairs, the first large-scale public benchmark for entity matching on sanctions and OSINT data. The dataset includes 755,540 expert-labeled pairs over 1 mil…