activity
20242026
most citedCoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents

Shidong Yang, Ziyu Ma, Tongwen Huang +5

Large language model (LLM) agents are trained with reinforcement learning (RL) for complex decision-making tasks. However, most RL-trained agents remain episodic and cannot accumul…

cs.CL2026

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

Hao Zhang, Yiming Hu, Yong Wang +3

Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the target model, enabling lossless…

cs.CL2026

SkillChain: Closing the Loop on Skill Evolution for Image-Based E-Commerce AI Assistants

Yimin Hu, Mengtao Xu, Hao Guo +3

Image-based AI assistants are now deployed at production scale on e-commerce platforms, where a single uploaded image can trigger fundamentally different user intents: product sear…

cs.CL20261 cited

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

Shidong Yang, Ziyu Ma, Tongwen Huang +3

Reinforcement learning for LLM agents is typically conducted on a static data distribution, which fails to adapt to the agent's evolving behavior and leads to poor coverage of comp…

cs.CL2024

Teaching Language Models to Self-Improve by Learning from Language Feedback

Chi Hu, Yimin Hu, Hang Cao +2

Aligning Large Language Models (LLMs) with human intentions and values is crucial yet challenging. Current methods primarily rely on human preferences, which are costly and insuffi…

cs.CL2024

Prior Constraints-based Reward Model Training for Aligning Large Language Models

Hang Zhou, Chenglong Wang, Yimin Hu +3

Reinforcement learning with human feedback for aligning large language models (LLMs) trains a reward model typically using ranking loss with comparison pairs.However, the training…