32 citations · 49 across the 15 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization
Dongze Hao, Zhiwei Jin, Chen Chen +1
On-policy distillation (OPD) has recently emerged as an important post-training paradigm. By using a stronger teacher model to provide dense, fine-grained supervision for sampled t…
cs.LG2026
fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Optimization (GRPO) serving as the…