1 paper
Evžen Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt
A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning h…