2 papers
cs.AI2026
Interpreting Language Model Hidden States at Scale
Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson +3
Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network.…
cs.LG2025
On Logical Extrapolation for Mazes with Recurrent and Implicit Networks
Brandon Knutson, Amandin Chyba Rabeendran, Michael Ivanitskiy +4
Recent work suggests that certain neural network architectures -- particularly recurrent neural networks (RNNs) and implicit neural networks (INNs) -- are capable of logical extrap…