most citedCoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

1 citations · 1 across the 21 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

Xucong Wang, Ziyu Ma, Yong Wang +5

Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…

cs.AI2026

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Xucong Wang, Ziyu Ma, Shidong Yang +4

Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…

cs.AI2026

Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution

Feng Xiong, Zengbin Wang, Yong Wang +5

Self-evolving agents present a promising path toward continual adaptation by distilling task interactions into reusable knowledge artifacts. In practice, this paradigm remains hind…

cs.AI2026

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Ziyu Ma, Shidong Yang, Yuxiang Ji +5

Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar w…

cs.AI2026

Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models

Shidong Yang, Tongwen Huang, Hao Wen +3

Multimodal reward models are crucial for aligning multimodal large language models with human preferences. Recent works have incorporated reasoning capabilities into these models,…

cs.AI2026

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

Yanqi Dai, Yuxiang Ji, Xiao Zhang +3

Reinforcement Learning with Verifiable Rewards (RLVR) offers a robust mechanism for enhancing mathematical reasoning in large models. However, we identify a systematic lack of emph…