2 papers
cs.AI2026
RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
Runyu Wang, Bo Liu, Xiaxin Zhang +6
Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estimates or compu…
cs.CL2025
LoKI: Low-damage Knowledge Implanting of Large Language Models
Runyu Wang, Peng Ping, Zhengyu Guo +4
Fine-tuning adapts pretrained models for specific tasks but poses the risk of catastrophic forgetting (CF), where critical knowledge from pretraining is overwritten. To address the…