1 paper · 1 filter
Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson +3
Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network.…