4 papers
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
Alan Sun, Mariya Toneva
Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpre…
General Recurrence Multidimensional Zeckendorf Representations
Jiarui Cheng, Steven J. Miller, Sebastian Rodriguez-Labastida +3
We present a multidimensional generalization of Zeckendorf's Theorem (any positive integer can be written uniquely as a sum of non-adjacent Fibonacci numbers) to a large family of…
Circuit Stability Characterizes Language Model Generalization
Alan Sun
Extensively evaluating the capabilities of (large) language models is difficult. Rapid development of state-of-the-art models induce benchmark saturation, while creating more chall…
Algorithmic Phase Transitions in Language Models: A Mechanistic Case Study of Arithmetic
Alan Sun, Ethan Sun, Warren Shepard
Zero-shot capabilities of large language models make them powerful tools for solving a range of tasks without explicit training. It remains unclear, however, how these models achie…