1 citations · 1 across the 8 of their papers we have counts for
7 papers · 1 filter
SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents
Shidong Yang, Ziyu Ma, Tongwen Huang +5
Large language model (LLM) agents are trained with reinforcement learning (RL) for complex decision-making tasks. However, most RL-trained agents remain episodic and cannot accumul…
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding
Hao Zhang, Yiming Hu, Yong Wang +3
Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the target model, enabling lossless…
SkillChain: Closing the Loop on Skill Evolution for Image-Based E-Commerce AI Assistants
Yimin Hu, Mengtao Xu, Hao Guo +3
Image-based AI assistants are now deployed at production scale on e-commerce platforms, where a single uploaded image can trigger fundamentally different user intents: product sear…
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
Shidong Yang, Ziyu Ma, Tongwen Huang +3
Reinforcement learning for LLM agents is typically conducted on a static data distribution, which fails to adapt to the agent's evolving behavior and leads to poor coverage of comp…
Teaching Language Models to Self-Improve by Learning from Language Feedback
Chi Hu, Yimin Hu, Hang Cao +2
Aligning Large Language Models (LLMs) with human intentions and values is crucial yet challenging. Current methods primarily rely on human preferences, which are costly and insuffi…
Prior Constraints-based Reward Model Training for Aligning Large Language Models
Hang Zhou, Chenglong Wang, Yimin Hu +3
Reinforcement learning with human feedback for aligning large language models (LLMs) trains a reward model typically using ranking loss with comparison pairs.However, the training…