Showing 2025Show all
3 papers · 1 filter
cs.LG2025
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Yiming Tang, Harshvardhan Saini, Zhaoqian Yao +6
As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important fo…
cs.DS2025
MagnifierSketch: Quantile Estimation Centered at One Point
Jiarui Guo, Qiushi Lyu, Yuhan Wu +6
In this paper, we take into consideration quantile estimation in data stream models, where every item in the data stream is a key-value pair. Researchers sometimes aim to estimate…
cs.CL2025
How Syntax Specialization Emerges in Language Models
Xufeng Duan, Zhaoqian Yao, Yunhao Zhang +2
Large language models (LLMs) have been found to develop surprising internal specializations: Individual neurons, attention heads, and circuits become selectively sensitive to synta…