collaborators

5 papers

cs.AI2026

SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees

Tianyi Hu, Qingxu Fu, Yanxi Chen +2

Reinforcement learning (RL) has emerged as the predominant paradigm for training large language model (LLM)-based AI agents. However, existing backbone RL algorithms lack verified…

cs.AI2025

CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL

Shinji Mai, Yunpeng Zhai, Ziqian Chen +5

Large language model based agents are increasingly deployed in complex, tool augmented environments. While reinforcement learning provides a principled mechanism for such agents to…

cs.LG2025

AgentEvolver: Towards Efficient Self-Evolving Agent System

Yunpeng Zhai, Shuchang Tao, Cheng Chen +10

Autonomous agents powered by large language models (LLMs) have the potential to significantly enhance human productivity by reasoning, using tools, and executing complex tasks in d…

cs.CL2025

MicroRemed: Benchmarking LLMs in Microservices Remediation

Lingzhe Zhang, Yunpeng Zhai, Tong Jia +6

Large Language Models (LLMs) integrated with agent-based reasoning frameworks have recently shown strong potential for autonomous decision-making and system-level operations. One p…

cs.LG2025

Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling

Lipeng Xie, Sen Huang, Zhuo Zhang +9

Conventional reward modeling relies on gradient descent over neural weights, creating opaque, data-hungry "black boxes." We propose a paradigm shift from implicit to explicit rewar…