Active Learning of Model Discrepancy with Bayesian Experimental Design
arXiv:2502.05372 · doi:10.1016/j.cma.2025.118198
Abstract
Digital twins have been actively explored in many engineering applications, such as manufacturing and autonomous systems. However, model discrepancy is ubiquitous in most digital twin models and has significant impacts on the performance of using those models. In recent years, data-driven modeling techniques have been demonstrated promising in characterizing the model discrepancy in existing models, while the training data for the learning of model discrepancy is often obtained in an empirical way and an active approach of gathering informative data can potentially benefit the learning of model discrepancy. On the other hand, Bayesian experimental design (BED) provides a systematic approach to gathering the most informative data, but its performance is often negatively impacted by the model discrepancy. In this work, we build on sequential BED and propose an efficient approach to iteratively learn the model discrepancy based on the data from the BED. The performance of the proposed method is validated by a classical numerical example governed by a convection-diffusion equation, for which full BED is still feasible. The proposed method is then further studied in the same numerical example with a high-dimensional model discrepancy, which serves as a demonstration for the scenarios where full BED is not practical anymore. An ensemble-based approximation of information gain is further utilized to assess the data informativeness and to enhance learning model discrepancy. The results show that the proposed method is efficient and robust to the active learning of high-dimensional model discrepancy, using data suggested by the sequential BED. We also demonstrate that the proposed method is compatible with both classical numerical solvers and modern auto-differentiable solvers.
38 pages, 14 figures
References in corpus (20)
- DeepONet: Learning nonlinear operators for identifying differential equations based on the universal approximation theorem of operators
- Machine Learning for Fluid Mechanics
- Turbulence Modeling in the Age of Data
- A Physics Informed Machine Learning Approach for Reconstructing Reynolds Stress Modeling Discrepancies Based on DNS Data
- A Comprehensive Physics-Informed Machine Learning Framework for Predictive Turbulence Modeling
- Physics-Informed Machine Learning Approach for Augmenting Turbulence Models: A Comprehensive Framework
- Simulation-based optimal Bayesian experimental design for nonlinear systems
- The Ensemble Kalman Filter for Inverse Problems
- Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks
- Gradient-based stochastic optimization methods in Bayesian experimental design
- Bayesian Sequential Optimal Experimental Design for Nonlinear Models Using Policy Gradient Reinforcement Learning
- Sequential Bayesian optimal experimental design via approximate dynamic programming
- Learning High-Dimensional Parametric Maps via Reduced Basis Adaptive Residual Networks
- Sequential Bayesian experimental design for calibration of expensive simulation models
- A Causality-Based Learning Approach for Discovering the Underlying Dynamics of Complex Systems from Partial Observations with Stochastic Parameterization
- Neural Dynamical Operator: Continuous Spatial-Temporal Model with Gradient-Based and Derivative-Free Optimization Methods
- CEBoosting: Online Sparse Identification of Dynamical Systems with Regime Switching by Causation Entropy Boosting
- CGNSDE: Conditional Gaussian Neural Stochastic Differential Equation for Modeling Complex Systems and Data Assimilation
- CGKN: A Deep Learning Framework for Modeling Complex Dynamical Systems and Efficient Data Assimilation
- Modeling Partially Observed Nonlinear Dynamical Systems and Efficient Data Assimilation via Discrete-Time Conditional Gaussian Koopman Network