1 paper · 1 filter
Tianyuan Shi, Canbin Huang, Bei Li +4
Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather th…