3 papers
cs.AI2026
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
Ruishan Fang, Siyuan Lu, Chenyi Zhuang +1
Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe that the gradient signal in GRPO concentrates on tasks with the…
cs.AI2026
Don't Just Fine-tune the Agent, Tune the Environment
Siyuan Lu, Zechuan Wang, Hongxuan Zhang +5
Large Language Model (LLM) agents show great promise for complex, multi-turn tool-use tasks, but their development is often hampered by the extreme scarcity of high-quality trainin…
cs.AI2025
AWorld: Orchestrating the Training Recipe for Agentic AI
Chengyue Yu, Siyuan Lu, Chenyi Zhuang +14
The learning from practice paradigm is crucial for developing capable Agentic AI systems, yet it is severely hampered by inefficient experience generation, a bottleneck especially…