1 paper
David Chanin, Anthony Hunter, Oana-Maria Camburu
Transformer language models (LMs) have been shown to represent concepts as directions in the latent space of hidden activations. However, for any human-interpretable concept, how c…