1 paper · 1 filter
Cedegao E. Zhang, Cédric Colas, Gabriel Poesia +2
Reasoning models (RMs), language models (LMs) trained with reinforcement learning to produce long-form natural language reasoning, have been remarkably successful, but they still r…