2 papers
cs.CL2026
There Is More to Refusal in Large Language Models than a Single Direction
Faaiz Joad, Majd Hawasly, Sabri Boughorbel +2
Prior work argues that refusal in large language models is mediated by a single activation-space direction, enabling effective steering and ablation. We show that this account is i…
cs.AI2026
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
Enes Altinisik, Masoomali Fatehkia, Fatih Deniz +4
Factual hallucination remains a central challenge for large language models (LLMs). Existing mitigation approaches primarily rely on either external post-hoc verification or mappin…