collaborators

9 papers

cs.LG2026

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

Ximo Zhu, Ruiqi Liu, Rong Wang +8

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local conf…

cs.LG2026

On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

Gengsheng Li, Mao Zheng, Mingyang Song +8

Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large…

cs.AI2026

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Zhengbo Zhang, Changtao Miao, Jinbo Su +10

Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex…

cs.LG2026

Dynamic resource matching in manufacturing using deep reinforcement learning

Saunak Kumar Panda, Yisha Xiang, Ruiqi Liu

Matching plays an important role in the logical allocation of resources across a wide range of industries. The benefits of matching have been increasingly recognized in manufacturi…

stat.ML2026

Online Statistical Inference of Constant Sample-averaged Q-Learning

Saunak Kumar Panda, Tong Li, Ruiqi Liu +1

Reinforcement learning algorithms have been widely used for decision-making tasks in various domains. However, the performance of these algorithms can be impacted by high variance…

cs.LG2026

R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training

Gengsheng Li, Jinghan He, Shijie Wang +7

Self-play bootstraps LLM reasoning through an iterative Challenger-Solver loop: the Challenger is trained to generate questions that target the Solver's capabilities, and the Solve…