1 citations · 1 across the 8 of their papers we have counts for
1 paper · 1 filter
Arne Vanhoyweghen, Brecht Verbeken, Andres Algaba +1
Fine-tuning Large Language Models (LLMs) with reinforcement learning to produce an explicit Chain-of-Thought (CoT) before answering produces models that consistently raise overall…