3 papers
cs.AI2026
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
Dekun Yang
A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-actio…
cs.CL2026
Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring
Dekun Yang
Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localizatio…
cs.AI2026
Calibrated Selective Fact-Checking via Evidence Chain Evaluation
Dekun Yang
Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts eve…