1 paper
Daniel A. Herrmann, Benjamin A. Levinstein
We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The co…