3 citations · 3 across the 7 of their papers we have counts for
1 paper · 1 filter
Qinglin Ye, Zhiyuan Gu, Jingjie Xia +6
Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues…