collaborators

8 papers

cs.LG2026

OPRD: On-Policy Representation Distillation

Shenzhi Yang, Guangcheng Zhu, Bowen Song +8

On-policy distillation (OPD) supervises the student exclusively in the output space by matching next-token distributions. This paradigm suffers from two limitations: (i) a high-var…

cs.LG2026

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling

Guangcheng Zhu, Shenzhi Yang, Haobo Wang +9

Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation cost…

cs.LG2026

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

Guangcheng Zhu, Shenzhi Yang, Haobo Wang +7

Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this…

cs.CL2026

GAPD: Gold-Action Policy Distillation for Agentic Reinforcement Learning in Knowledge Base Question Answering

Xin Sun, Jianan Xie, Zhongqi Chen +6

Reinforcement learning (RL) is a natural fit for agentic knowledge base question answering (KBQA), where a model must issue executable actions, observe knowledge-base feedback, and…

cs.LG2026

Can LLMs Learn to Reason Robustly under Noisy Supervision?

Shenzhi Yang, Guangcheng Zhu, Bowen Song +7

Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels du…

cs.CV2026

Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning

Ruoshuang Du, Xin Sun, Qiang Liu +4

Visual Question Answering systems face reliability issues due to hallucinations, where models generate answers misaligned with visual input or factual knowledge. While Retrieval Au…