DeepSeek-V3 Technical Report
arXiv:2412.19437
Abstract
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.
Cited by in corpus (20)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Large language models for automated scholarly paper review: A survey
- Reasoning Beyond Limits: Advances and Open Problems for LLMs
- Streamlining evidence based clinical recommendations with large language models
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Multi-step retrieval and reasoning improves radiology question answering with large language models
- Comparative Evaluation of ChatGPT and DeepSeek Across Key NLP Tasks: Strengths, Weaknesses, and Domain-Specific Performance
- Large Language Models for Depression Recognition in Spoken Language Integrating Psychological Knowledge
- RepairBench: Leaderboard of Frontier Models for Program Repair
- Instructor-Worker Large Language Model System for Policy Recommendation: a Case Study on Air Quality Analysis of the January 2025 Los Angeles Wildfires
- Adaptive Intelligence: leveraging insights from adaptive behavior in animals to build flexible AI systems
- EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- Perovskite-R1: a domain-specialized large language model for intelligent discovery of precursor additives and experimental design
- Mapping Diffuse Radio Sources Using TUNA: A Transformer-Based Deep Learning Approach
- A New DAPO Algorithm for Stock Trading
- SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
- Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
- A Roadmap for Tamed Interactions with Large Language Models
- English K_Quantization of LLMs Does Not Disproportionately Diminish Multilingual Performance
- Multi-agent Self-triage System with Medical Flowcharts