From the 1 of 15 linked papers with an AI index.
5 papers · 1 filter
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
Xucong Wang, Zhe Zhao, Liheng Yu +3
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
Xucong Wang, Ziyu Ma, Yong Wang +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Xucong Wang, Ziyu Ma, Shidong Yang +4
Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…
FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models
Xucong Wang, Pengkun Wang, Zhe Zhao +3
Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing m…
Spatiotemporal Causal Decoupling Model for Air Quality Forecasting
Jiaming Ma, Guanjun Wang, Sheng Huang +4
Due to the profound impact of air pollution on human health, livelihoods, and economic development, air quality forecasting is of paramount significance. Initially, we employ the c…