collaborators

10 papers

cs.LG2026

Verifier-Induced Support Reshaping in On-Policy Optimization

Shaohang Wei, Zikun Su, Feifan Song +4

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sa…

cs.CL2026

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

Guangyue Peng, Zongchao Chen, Wen Luo +9

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but answer-visible generation can justify a pre-committed answer rather than derive…

cs.CL2026

Encode Errors: Representational Retrieval of In-Context Demonstrations for Multilingual Grammatical Error Correction

Guangyue Peng, Wei Li, Wen Luo +1

Grammatical Error Correction (GEC) involves detecting and correcting the wrong usage of grammar. While large language models (LLMs) with in-context learning (ICL) capabilities have…

cs.CL2026

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

Wen Luo, Guangyue Peng, Liang Wang +7

Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors compound across reasoning s…

cs.CL2026

Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations

Wen Luo, Guangyue Peng, Wei Li +8

Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal states encode rich signals of truthf…

cs.CL2025

Mitigating Overthinking through Reasoning Shaping

Feifan Song, Shaohang Wei, Bofei Gao +8

Large reasoning models (LRMs) boosted by Reinforcement Learning from Verifier Reward (RLVR) have shown great power in problem solving, yet they often cause overthinking: excessive,…