1 citations · 1 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
A Controlled Study of Decoding-Time Truthfulness Methods on Instruction-Tuned LLMs
Ao Sun
Decoding-time truthfulness methods -- layer-contrast decoding, inference-time intervention, and learned logit adapters -- have demonstrated 10-30 point gains on TruthfulQA when app…
cs.CL2025
CHAIR -- Classifier of Hallucination as Improver
Ao Sun
In this work, we introduce CHAIR (Classifier of Hallucination As ImproveR), a supervised framework for detecting hallucinations by analyzing internal logits from each layer of ever…