collaborators

13 papers

cs.LG2026

Verifier-Induced Support Reshaping in On-Policy Optimization

Shaohang Wei, Zikun Su, Feifan Song +4

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sa…

cs.LG2026

Experience Augmented Policy Optimization for LLM Reasoning

Jinda Lu, Kexin Huang, Junkang Wu +7

Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR method…

cs.LG2026

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

Wenhao Yu, Shaohang Wei, Jiahong Liu +5

Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dimensional: the ground-truth probability…

cs.LG2026

One-Way Policy Optimization for Self-Evolving LLMs

Shuo Yang, Jinda Lu, Kexin Huang +6

Reinforcement Learning with Verifiable Rewards (RLVR) has become a promising paradigm for scaling reasoning capabilities of Large Language Models (LLMs). However, the sparsity of b…

cs.CL2026

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

Wen Luo, Guangyue Peng, Liang Wang +7

Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors compound across reasoning s…

cs.CL2026

Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations

Wen Luo, Guangyue Peng, Wei Li +8

Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal states encode rich signals of truthf…