9 papers
ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
Ximo Zhu, Ruiqi Liu, Rong Wang +8
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local conf…
On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents
Gengsheng Li, Mao Zheng, Mingyang Song +8
Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large…
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Zhengbo Zhang, Changtao Miao, Jinbo Su +10
Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex…
Dynamic resource matching in manufacturing using deep reinforcement learning
Saunak Kumar Panda, Yisha Xiang, Ruiqi Liu
Matching plays an important role in the logical allocation of resources across a wide range of industries. The benefits of matching have been increasingly recognized in manufacturi…
Online Statistical Inference of Constant Sample-averaged Q-Learning
Saunak Kumar Panda, Tong Li, Ruiqi Liu +1
Reinforcement learning algorithms have been widely used for decision-making tasks in various domains. However, the performance of these algorithms can be impacted by high variance…
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
Gengsheng Li, Jinghan He, Shijie Wang +7
Self-play bootstraps LLM reasoning through an iterative Challenger-Solver loop: the Challenger is trained to generate questions that target the Solver's capabilities, and the Solve…