collaborators

5 papers

cs.LG2026

Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control

Qi Zhao, Guozheng Ma, Yilun Kong +9

Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many…

cs.AI2026

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations

Lirong Che, Yuzhe yang, Peiwen lin +3

Agent harness evolution improves frozen language-model agents by modifying the executable structures around them. We study this paradigm as a form of sample-efficient fast adaptati…

cs.AI2025

GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning

Jiaqi Wu, Qinlao Zhao, Zefeng Chen +4

Autonomous agents powered by large language models (LLMs) have shown impressive capabilities in tool manipulation for complex task-solving. However, existing paradigms such as ReAc…

cs.LG2025

Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer

Yilun Kong, Guozheng Ma, Qi Zhao +4

Despite recent advancements in offline multi-task reinforcement learning (MTRL) have harnessed the powerful capabilities of the Transformer architecture, most approaches focus on a…

cs.CR2025

Lifelong Safety Alignment for Language Models

Haoyu Wang, Zeyu Qin, Yifei Zhao +4

LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing…