5 papers
Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training
Hang Chen, Jiaying Zhu, Hongyang Chen +3
The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpreta…
Skill Path: Unveiling Language Skills from Circuit Graphs
Hang Chen, Jiaying Zhu, Xinyu Yang +1
Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffe…
CLUE: Conflict-guided Localization for LLM Unlearning Framework
Hang Chen, Jiaying Zhu, Xinyu Yang +1
The LLM unlearning aims to eliminate the influence of undesirable data without affecting causally unrelated information. This process typically involves using a forget set to remov…
Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
Hang Chen, Jiaying Zhu, Xinyu Yang +1
Circuit discovery has gradually become one of the prominent methods for mechanistic interpretability, and research on circuit completeness has also garnered increasing attention. M…
Quantifying Semantic Emergence in Language Models
Hang Chen, Xinyu Yang, Jiaying Zhu +1
Large language models (LLMs) are widely recognized for their exceptional capacity to capture semantics meaning. Yet, there remains no established metric to quantify this capability…