1 paper
Joongwon Kim, Anirudh Goyal, Liang Tan +3
We introduce ASTRO, the "Autoregressive Search-Taught Reasoner", a framework for training language models to reason like search algorithms, explicitly leveraging self-reflection, b…