Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Scalable Circuit Learning for Interpreting Large Language Models
Naiyu Yin, Dennis Wei, Tian Gao +3
A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior. However, raw neuro…
cs.LG2024
Identifying Sub-networks in Neural Networks via Functionally Similar Representations
Tian Gao, Amit Dhurandhar, Karthikeyan Natesan Ramamurthy +1
Providing human-understandable insights into the inner workings of neural networks is an important step toward achieving more explainable and trustworthy AI. Existing approaches to…