Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
Ge Yan, Tuomas Oikarinen, Tsui-Wei +1
Neuron identification is a popular tool in mechanistic interpretability, aiming to uncover the human-interpretable concepts represented by individual neurons in deep networks. Whil…
cs.AI2025
ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering
Ge Yan, Chung-En Sun, Linbo Liu +3
Large reasoning models achieve strong performance on diverse tasks by producing extended chains of thought. Self-reflection, the ability to review and revise prior reasoning steps,…