3 papers
cs.LG2025
CLUE: Conflict-guided Localization for LLM Unlearning Framework
Hang Chen, Jiaying Zhu, Xinyu Yang +1
The LLM unlearning aims to eliminate the influence of undesirable data without affecting causally unrelated information. This process typically involves using a forget set to remov…
cs.LG2025
Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
Hang Chen, Jiaying Zhu, Xinyu Yang +1
Circuit discovery has gradually become one of the prominent methods for mechanistic interpretability, and research on circuit completeness has also garnered increasing attention. M…
cs.CL2024
Skill Path: Unveiling Language Skills from Circuit Graphs
Hang Chen, Jiaying Zhu, Xinyu Yang +1
Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffe…