1 paper
Nils Grünefeld, Bertram Højer, Philipp Mondorf +5
Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain…