1 paper · 1 filter
Raphael Ronge, Markus Maier, Frederick Eberhardt
Recent work by Anthropic on Mechanistic interpretability claims to understand and control Large Language Models by extracting human-interpretable features from their neural activat…