Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
arXiv:2505.17923
Abstract
Implicit reasoning is the ability of a language model to solve multi-hop reasoning tasks in a single forward pass, without chain of thought. We investigate this capability using GPT2-style language models trained from scratch on controlled -hop reasoning datasets (). We show that while such models can indeed learn implicit -hop reasoning, the required training data grows exponentially in , and the required number of transformer layers grows linearly in . We offer a theoretical explanation for why this depth growth is necessary. We further find that the data requirement can be mitigated, but not eliminated, through curriculum learning.
Accepted at EMNLP 2025