2 papers
cs.AI2026
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
Xuanzhi Feng, Zhengyang Li, Zeyu Liu +6
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a fundamental optimization instability: uni…
cs.LG2026
Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
Jianxiong Zhang, Bing Guo, Yuming Jiang +3
Large reasoning models (LRMs) often generate long, seemingly coherent reasoning traces yet still produce incorrect answers, making hallucination detection challenging. Although tra…