activity
20242026
most citedDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

978 citations · 1.5k across the 15 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.CL2024268 cited

DeepSeek-V3 Technical Report

DeepSeek-AI, Aixin Liu, Bei Feng +195

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effec…

cs.DC2024

Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning

Wei An, Xiao Bi, Guanting Chen +49

The rapid progress in Deep Learning (DL) and Large Language Models (LLMs) has exponentially increased demands of computational power and bandwidth. This, combined with the high cos…

cs.SE202449 cited

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

DeepSeek-AI, Qihao Zhu, Daya Guo +37

We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, D…

cs.AI202447 cited

DeepSeek-VL: Towards Real-World Vision-Language Understanding

Haoyu Lu, Wen Liu, Bo Zhang +12

We present DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. Our approach is structured around three ke…

cs.CL202418 cited

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Damai Dai, Chengqi Deng, Chenggang Zhao +14

In the era of large language models, Mixture-of-Experts (MoE) is a promising architecture for managing computational costs when scaling up model parameters. However, conventional M…

cs.CL202495 cited

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

DeepSeek-AI, :, Xiao Bi +85

The rapid development of open-source large language models (LLMs) has been truly remarkable. However, the scaling law described in previous literature presents varying conclusions,…