Accuracy Can Lie: On the Impact of Surrogate Model in Configuration Tuning
arXiv:2501.01876 · doi:10.1109/TSE.2025.3525955
Abstract
To ease the expensive measurements during configuration tuning, it is natural to build a surrogate model as the replacement of the system, and thereby the configuration performance can be cheaply evaluated. Yet, a stereotype therein is that the higher the model accuracy, the better the tuning result would be. This "accuracy is all" belief drives our research community to build more and more accurate models and criticize a tuner for the inaccuracy of the model used. However, this practice raises some previously unaddressed questions, e.g., Do those somewhat small accuracy improvements reported in existing work really matter much to the tuners? What role does model accuracy play in the impact of tuning quality? To answer those related questions, we conduct one of the largest-scale empirical studies to date-running over the period of 13 months 24*7-that covers 10 models, 17 tuners, and 29 systems from the existing works while under four different commonly used metrics, leading to 13,612 cases of investigation. Surprisingly, our key findings reveal that the accuracy can lie: there are a considerable number of cases where higher accuracy actually leads to no improvement in the tuning outcomes (up to 58% cases under certain setting), or even worse, it can degrade the tuning quality (up to 24% cases under certain setting). We also discover that the chosen models in most proposed tuners are sub-optimal and that the required % of accuracy change to significantly improve tuning quality varies according to the range of model accuracy. Deriving from the fitness landscape analysis, we provide in-depth discussions of the rationale behind, offering several lessons learned as well as insights for future opportunities. Most importantly, this work poses a clear message to the community: we should take one step back from the natural "accuracy is all" belief for model-based configuration tuning.
This paper has been accepted by TSE
References in corpus (16)
- BestConfig: Tapping the Performance Potential of Systems via Automatic Configuration Tuning
- Towards Dynamic and Safe Configuration Tuning for Cloud Databases
- The Weights can be Harmful: Pareto Search versus Weighted Search in Multi-Objective Search-Based Software Engineering
- Multi-Objectivizing Software Configuration Tuning (for a single performance concern)
- LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications
- Do Performance Aspirations Matter for Guiding Software Configuration Tuning?
- Does Configuration Encoding Matter in Learning Software Performance? An Empirical Study on Encoding Schemes
- HINNPerf: Hierarchical Interaction Neural Network for Performance Prediction of Configurable Systems
- Are we Forgetting about Compositional Optimisers in Bayesian Optimisation?
- Adapting Multi-objectivized Software Configuration Tuning
- Planning Landscape Analysis for Self-Adaptive Systems
- Framework and Benchmarks for Combinatorial and Mixed-variable Bayesian Optimization
- Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental Evaluation
- LlamaTune: Sample-Efficient DBMS Configuration Tuning
- Distilled Lifelong Self-Adaptation for Configurable Systems
- SCOPE: Safe Exploration for Dynamic Computer Systems Optimization