Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models
Yuetian Lu, Ali Modarressi, Yihong Liu +1
Activation patching can identify a mixture-of-experts (MoE) block whose clean output restores a corrupted factual prediction. However, because the block output combines contributio…
cs.CL2026
Relational Linearity is a Predictor of Hallucinations
Yuetian Lu, Yihong Liu, Sebastian Gerstner +3
Hallucination is a central failure mode of language models (LMs). We focus on hallucinations in response to questions like: "Which instrument did Glenn Gould play?", but we ask the…