3 papers
cs.CL2026
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…
cs.CY2026
Postmortem avatars in grief therapy: Prospects, ethics, and governance
Joshua Hatherley, Sandrine R. Schiller, Iwan Williams +3
Postmortem avatars (PMAs) -- AI systems that simulate a deceased person by being fine-tuned on data they generated or that was generated about them -- have attracted growing schola…
cs.MA2025
Goal-Directedness is in the Eye of the Beholder
Nina Rajcic, Anders Søgaard
Our ability to predict the behavior of complex agents turns on the attribution of goals. Probing for goal-directed behavior comes in two flavors: Behavioral and mechanistic. The fo…