2 papers
cs.CL2026
There Is More to Refusal in Large Language Models than a Single Direction
Faaiz Joad, Majd Hawasly, Sabri Boughorbel +2
Prior work argues that refusal in large language models is mediated by a single direction, enabling steering and abliteration. We show that this account is incomplete: across diver…
cs.AI2026
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
Enes Altinisik, Masoomali Fatehkia, Fatih Deniz +4
Factual hallucination remains a central challenge for large language models (LLMs). Existing mitigation approaches primarily rely on either external post-hoc verification or mappin…