1 paper · 1 filter
Daniel A. Herrmann, Benjamin A. Levinstein
We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The co…