Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Interpreting Reinforcement Learning Agents with Susceptibilities
Chris Elliott, Einar Urdshals, David Quarel +1
Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We gener…
cs.LG2026
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
Chris Elliott, Daniel Murfet
These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The susceptibility of an observable …