5 papers
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
Kexin Huang, Haoming Meng, Junkang Wu +10
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models. While existing analyses identify that RLVR-ind…
CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
Shinji Mai, Yunpeng Zhai, Ziqian Chen +5
Large language model based agents are increasingly deployed in complex, tool augmented environments. While reinforcement learning provides a principled mechanism for such agents to…
AgentEvolver: Towards Efficient Self-Evolving Agent System
Yunpeng Zhai, Shuchang Tao, Cheng Chen +10
Autonomous agents powered by large language models (LLMs) have the potential to significantly enhance human productivity by reasoning, using tools, and executing complex tasks in d…
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Dawei Gao, Zitao Li, Yuexiang Xie +20
Driven by rapid advancements of Large Language Models (LLMs), agents are empowered to combine intrinsic knowledge with dynamic tool use, greatly enhancing their capacity to address…
Larger or Smaller Reward Margins to Select Preferences for Alignment?
Kexin Huang, Junkang Wu, Ziqian Chen +6
Preference learning is critical for aligning large language models (LLMs) with human values, with the quality of preference datasets playing a crucial role in this process. While e…