1 paper · 1 filter
Abhranil Chandra, Ayush Agrawal, Arian Hosseini +4
We present the surprising finding that a language model's reasoning capabilities can be improved by training on synthetic datasets of chain-of-thought (CoT) traces from more capabl…