Insights into resource utilization of code small language models serving with runtime engines and execution providers
arXiv:2412.15441 · doi:10.1016/j.jss.2025.112574
Abstract
The rapid growth of language models, particularly in code generation, requires substantial computational resources, raising concerns about energy consumption and environmental impact. Optimizing language models inference resource utilization is crucial, and Small Language Models (SLMs) offer a promising solution to reduce resource demands. Our goal is to analyze the impact of deep learning serving configurations, defined as combinations of runtime engines and execution providers, on resource utilization, in terms of energy consumption, execution time, and computing-resource utilization from the point of view of software engineers conducting inference in the context of code generation SLMs. We conducted a technology-oriented, multi-stage experimental pipeline using twelve code generation SLMs to investigate energy consumption, execution time, and computing-resource utilization across the configurations. Significant differences emerged across configurations. CUDA execution provider configurations outperformed CPU execution provider configurations in both energy consumption and execution time. Among the configurations, TORCH paired with CUDA demonstrated the greatest energy efficiency, achieving energy savings from 37.99% up to 89.16% compared to other serving configurations. Similarly, optimized runtime engines like ONNX with the CPU execution provider achieved from 8.98% up to 72.04% energy savings within CPU-based configurations. Also, TORCH paired with CUDA exhibited efficient computing-resource utilization. Serving configuration choice significantly impacts resource utilization. While further research is needed, we recommend the above configurations best suited to software engineers' requirements for enhancing serving resource utilization efficiency.
Accepted in Journal of Systems and Software (JSS). For its published version refer to the Journal of JSS
References in corpus (26)
- Scaling Laws for Neural Language Models
- Evaluating Large Language Models Trained on Code
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Compute and Energy Consumption Trends in Deep Learning Inference
- Artificial Intelligence Index Report 2024
- InCoder: A Generative Model for Code Infilling and Synthesis
- Textbooks Are All You Need
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- On the Effectiveness of Large Language Models in Domain-Specific Code Generation
- pCAMP: Performance Comparison of Machine Learning Packages on the Edges
- Efficient Training of Language Models to Fill in the Middle
- A Systematic Review of Green AI
- Greening Large Language Models of Code
- Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures
- A Survey on Efficient Inference for Large Language Models
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
- Model Compression and Efficient Inference for Large Language Models: A Survey
- On the Opportunities of Green Computing: A Survey
- Identifying architectural design decisions for achieving green ML serving
- InferBench: Understanding Deep Learning Inference Serving with an Automatic Benchmarking System
- A Survey of Small Language Models
- EnergiBridge: Empowering Software Sustainability through Cross-Platform Energy Measurement
- Energy Consumption of Automated Program Repair
- Towards green AI-based software systems: an architecture-centric approach (GAISSA)
- GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation