AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
arXiv:2408.05100 · doi:10.1145/3691620.3695017
Abstract
Performance testing aims at uncovering efficiency issues of software systems. In order to be both effective and practical, the design of a performance test must achieve a reasonable trade-off between result quality and testing time. This becomes particularly challenging in Java context, where the software undergoes a warm-up phase of execution, due to just-in-time compilation. During this phase, performance measurements are subject to severe fluctuations, which may adversely affect quality of performance test results. However, these approaches often provide suboptimal estimates of the warm-up phase, resulting in either insufficient or excessive warm-up iterations, which may degrade result quality or increase testing time. There is still a lack of consensus on how to properly address this problem. Here, we propose and study an AI-based framework to dynamically halt warm-up iterations at runtime. Specifically, our framework leverages recent advances in AI for Time Series Classification (TSC) to predict the end of the warm-up phase during test execution. We conduct experiments by training three different TSC models on half a million of measurement segments obtained from JMH microbenchmark executions. We find that our framework significantly improves the accuracy of the warm-up estimates provided by state-of-practice and state-of-the-art methods. This higher estimation accuracy results in a net improvement in either result quality or testing time for up to +35.3% of the microbenchmarks. Our study highlights that integrating AI to dynamically estimate the end of the warm-up phase can enhance the cost-effectiveness of Java performance testing.
Accepted for publication in The 39th IEEE/ACM International Conference on Automated Software Engineering (ASE '24)
References in corpus (9)
- Deep learning for time series classification: a review
- Optimal detection of changepoints with a linear computational cost
- Scalable and accurate deep learning for electronic health records
- ROCKET: Exceptionally fast and accurate time series classification using random convolutional kernels
- Virtual Machine Warmup Blows Hot and Cold
- Applications of shapelet transform to time series classification of earthquake, wind and wave data
- Multi-Objectivizing Software Configuration Tuning (for a single performance concern)
- Towards effective assessment of steady state performance in Java software: Are we there yet?
- Evaluating Search-Based Software Microbenchmark Prioritization