activity
20242026
collaborators

6 papers

cs.LG2026

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

Xian Qi Loye, Qinglin Su, Zhexin Zhang +5

The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text generation. Existing alignment…

cs.AI2026

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

Xujun Li, Kehan Zheng, Mingyuan Zhao +7

Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving strategies or parametric lea…

cs.AI2026

SkillEvolver: Skill Learning as a Meta-Skill

Genrui Zhang, Erle Zhu, Jinfeng Zhou +2

Agent skills today are static artifact: authored once -- by human curation or one-shot generation from parametric knowledge -- and then consumed unchanged, with no mechanism to imp…

cs.LG2025

Trust-Region Adaptive Policy Optimization

Mingyu Su, Jian Guan, Yuxian Gu +2

Post-training methods, especially Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), play an important role in improving large language models' (LLMs) complex reasoning…

cs.CL2024

AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback

Jian Guan, Wei Wu, Zujie Wen +3

The notable success of large language models (LLMs) has sparked an upsurge in building language agents to complete various complex tasks. We present AMOR, an agent framework based…

cs.CL2024

Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning

Jiaxin Wen, Jian Guan, Hongning Wang +2

Despite the remarkable success of large language models (LLMs) on traditional natural language processing tasks, their planning ability remains a critical bottleneck in tackling co…