4 papers
Miles v0.1: Production-Level Post-Training
RadixArk, :, Tom Chen +11
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-lear…
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
Ruishan Fang, Siyuan Lu, Chenyi Zhuang +1
Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe that the gradient signal in GRPO concentrates on tasks with the…
Don't Just Fine-tune the Agent, Tune the Environment
Siyuan Lu, Zechuan Wang, Hongxuan Zhang +5
Large Language Model (LLM) agents show great promise for complex, multi-turn tool-use tasks, but their development is often hampered by the extreme scarcity of high-quality trainin…
AWorld: Orchestrating the Training Recipe for Agentic AI
Chengyue Yu, Siyuan Lu, Chenyi Zhuang +14
The learning from practice paradigm is crucial for developing capable Agentic AI systems, yet it is severely hampered by inefficient experience generation, a bottleneck especially…