4 papers
SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
Bingqing Jiang, Guoxi Zhang, Jasper Wang +7
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. Howe…
From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models
Bingqing Jiang, Li Luo, Zichao Yu +3
On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models ar…
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
Zichao Yu, Chengzhi Yu, Shengze Xu +4
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repet…
Relative Score Policy Optimization for Diffusion Language Models
Zichao Yu, Shengze Xu, Bingqing Jiang +2
Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability requires effective post-training. R…