4 papers
A Mechanistic Analysis of Looped Reasoning Language Models
Hugh Blayney, Ãlvaro Arroyo, Johan Obando-Ceron +4
Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM's layers in the latent di…
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Milad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad +4
Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment…
VinePPO: Refining Credit Assignment in RL Training of LLMs
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance +4
Large language models (LLMs) are increasingly applied to complex reasoning tasks that require executing several complex steps before receiving any reward. Properly assigning credit…
Not All LLM Reasoners Are Created Equal
Arian Hosseini, Alessandro Sordoni, Daniel Toyama +2
We study the depth of grade-school math (GSM) problem-solving capabilities of LLMs. To this end, we evaluate their performance on pairs of existing math word problems together so t…