collaborators

7 papers

cs.LG2026

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Leyi Pan, Shuchang Tao, Yunpeng Zhai +5

On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distribution it produces under privi…

cs.AI2026

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning

Qingxu Fu, Boyin Liu, Shuchang Tao +5

Training reinforcement learning (RL) policies for large language model (LLM) agents requires optimizing multi-turn trajectories that interact with external environments. Existing t…

cs.CL2026

d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models

Leyi Pan, Shuchang Tao, Yunpeng Zhai +8

Reinforcement learning (RL) is pivotal for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, existing dLLM policy optimization methods suffe…

cs.AI2025

CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL

Shinji Mai, Yunpeng Zhai, Ziqian Chen +5

Large language model based agents are increasingly deployed in complex, tool augmented environments. While reinforcement learning provides a principled mechanism for such agents to…

cs.LG2025

AgentEvolver: Towards Efficient Self-Evolving Agent System

Yunpeng Zhai, Shuchang Tao, Cheng Chen +10

Autonomous agents powered by large language models (LLMs) have the potential to significantly enhance human productivity by reasoning, using tools, and executing complex tasks in d…

cs.CL2025

Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models

Leyi Pan, Zheyu Fu, Yunpeng Zhai +9

Omni-modal Large Language Models (OLLMs) that integrate visual, auditory, and textual processing face severe safety risks. They exhibit fragile defenses against audio-visual joint…