4 papers
Benign interpolation and Occam's razor
Tom F. Sterkenburg, Daniel A. Herrmann, Jan-Willem Romeijn
Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation. This phenomenon cannot be accounted…
Radical AI Interpretability
Daniel A. Herrmann, Benjamin A. Levinstein
We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The co…
A Decision-Theoretic Approach for Managing Misalignment
Daniel A. Herrmann, Abinav Chari, Isabelle Qian +2
When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determi…
Standards for Belief Representations in LLMs
Daniel A. Herrmann, Benjamin A. Levinstein
As large language models (LLMs) continue to demonstrate remarkable abilities across various domains, computer scientists are developing methods to understand their cognitive proces…