collaborators

14 papers

cs.AI2026

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Zhanming Shen, Jintao Tong, Shaotian Yan +9

On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-lev…

cs.LG2026

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling

Guangcheng Zhu, Shenzhi Yang, Haobo Wang +9

Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation cost…

cs.LG2026

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

Guangcheng Zhu, Shenzhi Yang, Haobo Wang +7

Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this…

cs.LG2026

Can LLMs Learn to Reason Robustly under Noisy Supervision?

Shenzhi Yang, Guangcheng Zhu, Bowen Song +7

Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels du…

cs.LG2026

FastBUS: A Fast Bayesian Framework for Unified Weakly-Supervised Learning

Ziquan Wang, Haobo Wang, Ke Chen +2

Machine Learning often involves various imprecise labels, leading to diverse weakly supervised settings. While recent methods aim for universal handling, they usually suffer from c…

cs.CL2026

Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority

Zhanming Shen, Zeyu Qin, Jiaqi Hu +7

The transition from fitting empirical data to achieving true human utility is fundamentally constrained by a granularity mismatch, where fine-grained autoregressive generation is o…