1 paper · 1 filter
Jungseob Lee, Seungyoon Lee, Suhyune Son +4
A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer…