2 papers
cs.LG2026
LLMs Can Annotate Attribution Graphs
Ameen Patel, Max Zhang, Nathan Hu
Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or…
cs.AI2025
Measuring Sparse Autoencoder Feature Sensitivity
Claire Tian, Katherine Tian, Nathan Hu
Sparse Autoencoder (SAE) features have become essential tools for mechanistic interpretability research. SAE features are typically characterized by examining their activating exam…