1 paper
Fangcong Yin, Zeyu Leo Liu, Liu Leqi +2
A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoning problems, but such annotated data is…