activity
20242026
collaborators

6 papers

cs.CL2026

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training

Hang Chen, Jiaying Zhu, Hongyang Chen +3

The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpreta…

cs.CL2025

Skill Path: Unveiling Language Skills from Circuit Graphs

Hang Chen, Jiaying Zhu, Xinyu Yang +1

Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffe…

cs.LG2025

CLUE: Conflict-guided Localization for LLM Unlearning Framework

Hang Chen, Jiaying Zhu, Xinyu Yang +1

The LLM unlearning aims to eliminate the influence of undesirable data without affecting causally unrelated information. This process typically involves using a forget set to remov…

cs.LG2025

Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates

Hang Chen, Jiaying Zhu, Xinyu Yang +1

Circuit discovery has gradually become one of the prominent methods for mechanistic interpretability, and research on circuit completeness has also garnered increasing attention. M…

cs.LG2024

SSL Framework for Causal Inconsistency between Structures and Representations

Hang Chen, Xinyu Yang, Keqing Du +1

The cross-pollination between causal discovery and deep learning has led to increasingly extensive interactions. It results in a large number of deep learning data types (such as i…

cs.CL2024

Quantifying Semantic Emergence in Language Models

Hang Chen, Xinyu Yang, Jiaying Zhu +1

Large language models (LLMs) are widely recognized for their exceptional capacity to capture semantics meaning. Yet, there remains no established metric to quantify this capability…