works on

From the 1 of 14 linked papers with an AI index.

collaborators

14 papers

cs.AI2026

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

Xucong Wang, Zhe Zhao, Liheng Yu +3

Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…

cs.CV2026

DAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image Classification

Pengkun Wang, Weijia Cao, Ning Wang +1

The paper proposes DAPGNet, a graph diffusion network that incorporates physical priors from contiguous spectral bands to improve hyperspectral image classification, using adaptive…

cs.AI2026

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

Xucong Wang, Ziyu Ma, Yong Wang +5

Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…

cs.LG2026

APPO: Agentic Procedural Policy Optimization

Xucong Wang, Ziyu Ma, Yong Wang +5

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…

cs.AI2026

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Xucong Wang, Ziyu Ma, Shidong Yang +4

Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…

cs.AI2026

FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models

Xucong Wang, Pengkun Wang, Zhe Zhao +3

Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing m…