Leveraging Remote Traffic Data for Local Air Pollutant Estimation: A Scenario-Based Machine Learning Study Across London Monitoring Sites
arXiv:2608.23219 · doi:10.3390/atmos17080806
Abstract
Vehicular traffic is a major source of air pollution; however, the contribution of remotely acquired traffic information to local machine-learning (ML) air-pollution models remains insufficiently characterised. This study evaluates four interpretable tree-based ML models (Random Forest, Extra Trees, LightGBM, and XGBoost) under six predictor scenarios combining progressively larger predictor sets, ranging from remotely acquired traffic, meteorological, and temporal variables alone to the inclusion of measurements from one and four neighbouring monitoring stations, to estimate NO, PM, PM, and O concentrations across several sites in London. ML model performance was compared with a ridge linear regression model as a baseline, with spatial interpolation methods and with a cross-site validation experiment. When modelling without data from neighbouring stations, the RMSE for NO ranged from 9.73 to 11.66 g/m without traffic information, compared with 8.72 to 11.52 g/m when traffic information was included. Additionally, for NO, SHAP analyses indicate that traffic-related variables can contribute at levels comparable to pollutant measurements from neighbouring monitoring stations in traffic-dominated~environments.
Accepted for publication in Atmosphere