Maximizing information from chemical engineering data sets: Applications to machine learning
arXiv:2201.10035 · doi:10.1016/j.ces.2022.117469
Abstract
It is well-documented how artificial intelligence can have (and already is having) a big impact on chemical engineering. But classical machine learning approaches may be weak for many chemical engineering applications. This review discusses how challenging data characteristics arise in chemical engineering applications. We identify four characteristics of data arising in chemical engineering applications that make applying classical artificial intelligence approaches difficult: (1) high variance, low volume data, (2) low variance, high volume data, (3) noisy/corrupt/missing data, and (4) restricted data with physics-based limitations. For each of these four data characteristics, we discuss applications where these data characteristics arise and show how current chemical engineering research is extending the fields of data science and machine learning to incorporate these challenges. Finally, we identify several challenges for future research.
34 pages, 3 figures, 1 table
References in corpus (7)
- Big Data of Materials Science - Critical Role of the Descriptor
- Optimization under Uncertainty in the Era of Big Data and Deep Learning: When Machine Learning Meets Mathematical Programming
- Resolving transition metal chemical space: feature selection for machine learning and structure-property relationships
- Understanding molecular representations in machine learning: The role of uniqueness and target similarity
- Computer-aided molecular design: An introduction and review of tools, applications, and solution techniques
- The ALAMO approach to machine learning
- Multi-Objective Constrained Optimization for Energy Applications via Tree Ensembles