LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
arXiv:2504.03665 · doi:10.1007/978-3-032-07612-0_48
Abstract
Large Language Models (LLMs), such as GPT-4 and DeepSeek, have been applied to a wide range of domains in software engineering. However, their potential in the context of High-Performance Computing (HPC) much remains to be explored. This paper evaluates how well DeepSeek, a recent LLM, performs in generating a set of HPC benchmark codes: a conjugate gradient solver, the parallel heat equation, parallel matrix multiplication, DGEMM, and the STREAM triad operation. We analyze DeepSeek's code generation capabilities for traditional HPC languages like Cpp, Fortran, Julia and Python. The evaluation includes testing for code correctness, performance, and scaling across different configurations and matrix sizes. We also provide a detailed comparison between DeepSeek and another widely used tool: GPT-4. Our results demonstrate that while DeepSeek generates functional code for HPC tasks, it lags behind GPT-4, in terms of scalability and execution efficiency of the generated code.
9 pages, 2 figures, 3 tables, conference
References in corpus (7)
- HPC-GPT: Integrating Large Language Model for High-Performance Computing
- HPC-Coder: Modeling Parallel Programs using Large Language Models
- Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation
- LM4HPC: Towards Effective Language Model Application in High-Performance Computing
- Benchmarking the Parallel 1D Heat Equation Solver in Chapel, Charm++, C++, HPX, Go, Julia, Python, Rust, Swift, and Java
- LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
- Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust