A sampling method based on highest density regions: Applications to surrogate models
arXiv:2509.10149 · doi:10.1016/j.strusafe.2026.102750
Abstract
This paper introduces a practical, one-shot design-of-experiments strategy for training surrogate models in the context of uncertainty quantification. Instead of drawing the experimental design according to the input distribution, which favours training points around its mode, we propose a heuristic method to sample uniformly within a highest density region (HDR) of the input random vector, aiming at a more homogeneous approximation accuracy over the whole domain. The approach is assessed through four error metrics: two task-agnostic accuracy measures (leave-one-out and relative mean square error) and two task-oriented measures quantifying the recovery of reference global sensitivity indices and of the probability of failure. Across the benchmark, the HDR-based designs are comparable to, or more accurate than, the distribution-based designs for most surrogate-task combinations, with the clearest gains in predictive accuracy and reliability and essentially no effect on the recovery of sensitivity indices. The method operates in a black-box context and is compatible with existing uncertainty quantification frameworks for low-dimensional and moderately correlated inputs, the curse of dimensionality affecting the sample generation rather than the design strategy itself. It is most useful whenever a single surrogate is reused across several downstream UQ tasks.
31 pages, 16 figures