Publications (9)
Quantifying Local Specialization in Deep Neural Networks
Shlomi Hod, Daniel Filan, Stephen Casper +2
A neural network is locally specialized to the extent that parts of its computational graph (i.e. structure) can be abstractly represented as performing some comprehensible sub-tas…
Clusterability in Neural Networks
Daniel Filan, Stephen Casper, Shlomi Hod +3
The learned weights of a neural network have often been considered devoid of scrutable internal structure. In this paper, however, we look for structure in the form of clusterabili…
Constrained belief updates explain geometric structures in transformer representations
Mateusz Piotrowski, Paul M. Riechers, Daniel Filan +1
What computational structures emerge in transformers trained on next-token prediction? In this work, we provide evidence that transformers implement constrained Bayesian belief upd…
Pruned Neural Networks are Surprisingly Modular
Daniel Filan, Shlomi Hod, Cody Wild +2
The learned weights of a neural network are often considered devoid of scrutable internal structure. To discern structure in these weights, we introduce a measurable notion of modu…
Loss Bounds and Time Complexity for Speed Priors
Daniel Filan, Marcus Hutter, Jan Leike
This paper establishes for the first time the predictive performance of speed priors and their computational complexity. A speed prior is essentially a probability distribution tha…
On the Impossibility of Supersized Machines
Ben Garfinkel, Miles Brundage, Daniel Filan +6
In recent years, a number of prominent computer scientists, along with academics in fields such as philosophy and physics, have lent credence to the notion that machines may one da…
Exploring Hierarchy-Aware Inverse Reinforcement Learning
Chris Cundy, Daniel Filan
We introduce a new generative model for human planning under the Bayesian Inverse Reinforcement Learning (BIRL) framework which takes into account the fact that humans often plan u…
What would it have looked like if it looked like I were in a superposition?
Daniel Filan, Joseph J. Hope
In this paper we address the question of whether it is possible to obtain evidence that we are in a superposition of different worlds, as suggested by the relative state interpreta…
Self-Modification of Policy and Utility Function in Rational Agents
Tom Everitt, Daniel Filan, Mayank Daswani +1
Any agent that is part of the environment it interacts with and has versatile actuators (such as arms and fingers), will in principle have the ability to self-modify -- for example…