1 paper
Terry Tong, Yu Feng, Surbhi Goel +1
For tool-augmented language models, comparing natural-language reasoning with code-execution pipelines is difficult because the comparison changes both the intermediate representat…