Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
Alan Sun, Mariya Toneva
Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpre…
cs.LG2024
Achieving Domain-Independent Certified Robustness via Knowledge Continuity
Alan Sun, Chiyu Ma, Kenneth Ge +1
We present knowledge continuity, a novel definition inspired by Lipschitz continuity which aims to certify the robustness of neural networks across input domains (such as continuou…