2 papers
cs.LG2026
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
Brandon Cui, Ximing Lu, Jaehun Jung +7
We tackle the question of how to scale more efficiently across the many, ever-growing stages of current LLM training pipelines. Our guiding intuition stems from the fact that the d…
cs.AI2026
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
Hao Zhang, Mingjie Liu, Shaokun Zhang +10
Multi-turn LLM agents are increasingly important for solving complex, interactive tasks, and reinforcement learning (RL) is a key ingredient for improving their long-horizon behavi…