collaborators

5 papers

cs.LG2026

On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation

Kexin Huang, Haoming Meng, Junkang Wu +10

Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models. While existing analyses identify that RLVR-ind…

cs.AI2025

CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL

Shinji Mai, Yunpeng Zhai, Ziqian Chen +5

Large language model based agents are increasingly deployed in complex, tool augmented environments. While reinforcement learning provides a principled mechanism for such agents to…

cs.LG2025

AgentEvolver: Towards Efficient Self-Evolving Agent System

Yunpeng Zhai, Shuchang Tao, Cheng Chen +10

Autonomous agents powered by large language models (LLMs) have the potential to significantly enhance human productivity by reasoning, using tools, and executing complex tasks in d…

cs.AI2025

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications

Dawei Gao, Zitao Li, Yuexiang Xie +20

Driven by rapid advancements of Large Language Models (LLMs), agents are empowered to combine intrinsic knowledge with dynamic tool use, greatly enhancing their capacity to address…

cs.LG2025

Larger or Smaller Reward Margins to Select Preferences for Alignment?

Kexin Huang, Junkang Wu, Ziqian Chen +6

Preference learning is critical for aligning large language models (LLMs) with human values, with the quality of preference datasets playing a crucial role in this process. While e…