From the 1 of 11 linked papers with an AI index.
11 papers
Simple-OPD: Demystifying Warm-up for On-policy Distillation
Tao Liu, Taiqiang Wu, Mao Zheng +5
On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage b…
Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA
Tianhang Yang, Yanze Zheng, Junjie Wang +3
The paper introduces KinRT, a kinematics‑supervised routing method that clusters action trajectories to guide expert selection in mixture‑of‑experts vision‑language agents for robo…
Decoupled Residual Quantization for Robust Semantic IDs in Recommendation
Xuesi Wang, Junjie Wang, Ziliang Wang +2
Semantic IDs represent items as shared discrete token sequences and have become a practical tool for recommendation and retrieval. Yet it remains difficult to tell why a tokenizer…
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
Xuewei Yang, Jiachen Yu, Jie Wu +3
Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated…
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
Jiachen Yu, Zhihao Xu, Junjie Wang +1
Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. Howeve…
OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing
Jianming Chen, Yawen Wang, Junjie Wang +3
Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, this capability introduces a new…