4 papers
Rational Sparse Autoencoder
Naiyu Yin, Yue Yu
Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder nonlinearities such as ReLU, JumpReLU, and…
Scalable Circuit Learning for Interpreting Large Language Models
Naiyu Yin, Dennis Wei, Tian Gao +3
A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior. However, raw neuro…
Learning Causal Graphs at Scale: A Foundation Model Approach
Naiyu Yin, Tian Gao, Yue Yu
Due to its human-interpretability and invariance properties, Directed Acyclic Graph (DAG) has been a foundational tool across various areas of AI research, leading to significant a…
Effective Causal Discovery under Identifiable Heteroscedastic Noise Model
Naiyu Yin, Tian Gao, Yue Yu +1
Capturing the underlying structural causal relations represented by Directed Acyclic Graphs (DAGs) has been a fundamental task in various AI disciplines. Causal DAG learning via th…