1 paper
Tobias Bersia, Tatiana Gaintseva
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hi…