4 papers
Iris: Climbing to the Search Frontier
Ziyuan Liu, Hengqi Liu, Zichuan Wang +6
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse…
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
Zhiwei Zhang, Zechen Sun, Fei Zhao +6
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supe…
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Yang Xiao, Yusong Sun, Haoyi Wu +7
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends,…
D-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
Zechen Sun, Zhiwei Zhang, Fei Zhao +7
Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollou…