22 citations · 55 across the 112 of their papers we have counts for
49 papers · 1 filter
Simple-OPD: Demystifying Warm-up for On-policy Distillation
Tao Liu, Taiqiang Wu, Mao Zheng +5
On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage b…
WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
Yuxin Meng, Yuhan Suo, Junjie Wang +9
Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page…
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
Yusong Zhao, Yuejin Xie, Youliang Yuan +4
Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable multimodal large language mod…
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
Xuewei Yang, Jiachen Yu, Jie Wu +3
Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated…
OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
Xinchen Zhang, Bowei Liu, Jiale Liu +7
Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In t…
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
Jiachen Yu, Zhihao Xu, Junjie Wang +1
Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. Howeve…