DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
arXiv:2501.12948 · doi:10.1038/s41586-025-09422-z
Abstract
General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.
References in corpus (7)
- Scaling Laws for Neural Language Models
- DeepSeek-V3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Instruction-Following Evaluation for Large Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Cited by in corpus (27)
- Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey
- Humanity's Last Exam
- Data-Centric Foundation Models in Computational Healthcare: A Survey
- InvDesFlow-AL: active learning-based workflow for inverse design of functional materials
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- Multi-step retrieval and reasoning improves radiology question answering with large language models
- AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
- Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
- Open and Sustainable AI: challenges, opportunities and the road ahead in the life sciences (October 2025 -- Version 2)
- Perovskite-R1: a domain-specialized large language model for intelligent discovery of precursor additives and experimental design
- GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols
- The asymmetric structure of the inner disc around HD 142527 A with VLTI/MATISSE
- Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs
- PhenoAssistant: A Conversational Multi-Agent AI System for Automated Plant Phenotyping
- The Open Syndrome Definition
- LLM-Confidence Reranker: A Training-Free Approach for Enhancing Retrieval-Augmented Generation Systems
- LLM-Based Multi-Agent Collaboration for Constrained Multi-Objective Container Placement
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- AutoSurrogate: An LLM-Driven Multi-Agent Framework for Autonomous Construction of Deep Learning Surrogate Models in Subsurface Flow
- KD-Judge: A Knowledge-Driven Automated Judge Framework for Functional Fitness Movements on Edge Devices
- Zoom In Disparities in Healthcare LLM Q&A
- Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
- Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
- MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation
- Large Language Models -- the Future of Fundamental Physics?
- Network Dynamics-Based Framework for Understanding Deep Neural Networks
- ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding