From the 1 of 14 linked papers with an AI index.
14 papers
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
Xucong Wang, Zhe Zhao, Liheng Yu +3
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…
DAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image Classification
Pengkun Wang, Weijia Cao, Ning Wang +1
The paper proposes DAPGNet, a graph diffusion network that incorporates physical priors from contiguous spectral bands to improve hyperspectral image classification, using adaptive…
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
Xucong Wang, Ziyu Ma, Yong Wang +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…
APPO: Agentic Procedural Policy Optimization
Xucong Wang, Ziyu Ma, Yong Wang +5
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Xucong Wang, Ziyu Ma, Shidong Yang +4
Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…
FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models
Xucong Wang, Pengkun Wang, Zhe Zhao +3
Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing m…