1 paper
Shardul Bansal, Seth Schilbe, Jarrod Barnes
Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes. Leading recipes (RLVR, group-relative RL, rejection-sampled re…