3 papers
cs.AI2026
Interpreting Language Model Hidden States at Scale
Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson +3
Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network.…
cs.LG2025
On Logical Extrapolation for Mazes with Recurrent and Implicit Networks
Brandon Knutson, Amandin Chyba Rabeendran, Michael Ivanitskiy +4
Recent work suggests that certain neural network architectures -- particularly recurrent neural networks (RNNs) and implicit neural networks (INNs) -- are capable of logical extrap…
cs.LG2025
Deep Model Merging: The Sister of Neural Network Interpretability -- A Survey
Arham Khan, Todd Nief, Nathaniel Hudson +6
We survey the model merging literature through the lens of loss landscape geometry to connect observations from empirical studies on model merging and loss landscape analysis to ph…