Search NASA⌕ Search

NASA NTRS · 20230009508

Reconstructing PM 2.5 Data Record for the Kathmandu Valley Using a Machine Learning Model

Abstract

This paper presents a method for reconstructing the historical hourly concentrations of particulate matter 2.5 (PM2.5) over the Kathmandu Valley from 1980 to the present. The method uses a machine learning model that is trained using PM2.5 readings from US Embassy (Phora Durbar) as a ground truth, and the meteorological data from Modern-Era Retrospective Analysis for Research and Applications v2 (MERRA2) as input. The Extreme Gradient Boosting (XGBoost) model acquires a credible 10-fold cross-validation (CV) score of ~83.4%, an r2-score of ~84%, a Root Mean Square Error (RMSE) of ~15.82 µg/m3, and a Mean Absolute Error (MAE) of ~10.27 µg/m3. Further demonstrating the model's applicability to years other than those for which truth values are unavailable, the multiple cross-test with an unseen data set offered r2-scores for 2018, 2019, and 2020 ranging from 56% to 67%. The model-predicted data agrees with true values and indicates that MERRA2 underestimates PM2.5 over the region. It strongly agrees with ground-based evidence showing substantially higher mass concentrations in the dry pre- and post-monsoon seasons than in the monsoon months. It also shows a strong anti-correlation between PM2.5 concentration and humidity. The results also demonstrate that none of the years fulfilled the annual mean air quality index (AQI) standards set by the World Health Organization (WHO).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Surendra Bhatta, Yuekui Yang. 2023-06-25. Reconstructing PM 2.5 Data Record for the Kathmandu Valley Using a Machine Learning Model. https://ntrs.nasa.gov/citations/20230009508

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Proactive Wildfire Management: A Remote Sensing and Multimodal CNN-MLP Architecture for Ignition Risk Forecasting

As the frequency and intensity of wildfires increase, with fire seasons now starting earlier and ending later than they have over the past decades, current monitoring systems, such as lookout towers and satellites, are hindered by cloud cover, low-resolution imagery, and static data gaps that fail to track vegetation moisture levels fast enough to catch rapid pre-ignition changes. This report proposes a Machine Learning-enabled Wildfire Ignition Prediction framework that combines satellite monitoring with dynamic and high-resolution remote sensing from Unmanned Aerial Vehicle (UAV) swarms. The method would use multispectral and thermal data from the Landsat program to create a baseline for vegetation health, calculating a two-band Enhanced Vegetation Index (EVI2) and the moisture content of the vegetation. These inputs will later be fused with microscale UAV weather data, including thermal hotspots found through thick canopies, hyperspectral chemical signatures of pre-visual combustion, and local weather streams. The multispectral satellite, multispectral Light Detection and Ranging (LiDAR), and thermal data would then be processed through a Convolutional Neural Network (CNN), alongside a Multilayer Perceptron (MLP) for the micro-weather telemetry. The outputs of these networks would be fused into a single feature representation and passed through a final prediction network to generate real-time ignition risk scores and hotspot alerts. Model performance would be assessed using standard classification metrics, including a Receiver Operating Characteristic - Area Under the Curve (ROC AUC) and F1 score. This system would allow first responders to identify high-risk zones and intervene before ignition occurs, improving emergency response time compared to current approaches.

machine learning↗

Enabling integrated AI control on DIII-D: a control system design with state-of-the-art experiments

We present the design and application of a general algorithm for Prediction And Control using MAchiNe learning (PACMAN) in DIII-D. Machine learning (ML)-based predictors and controllers have shown great promise in achieving regimes in which traditional controllers fail, such as tearing mode (TM) free scenarios, ELM-free scenarios and stable advanced tokamak conditions. The architecture presented here was deployed on DIII-D to facilitate the end-to-end implementation of advanced control experiments, from diagnostic processing to final actuation commands. This paper describes the detailed design of the algorithm and explains the motivation behind each design point. We also describe several successful ML control experiments in DIII-D using this algorithm, including a reinforcement learning controller targeting advanced non-inductive plasmas, a wide-pedestal quiescent H-mode ELM predictor, an Alfvén Eigenmode controller, a Model Predictive Control plasma profile controller and a state-machine TM predictor-controller. There is also discussion on guiding principles for real-time ML controller design and implementation.

machine learning↗

Protonation Dynamics of Confined Ethanol–Water Mixtures in H-ZSM-5 from Machine Learning-Driven Metadynamics

Zeolites are indispensable heterogeneous catalysts in industrial chemical processes, valued for their strong Brønsted acidity, well-defined microporous frameworks, and tunable pore structures. Their catalytic activity arises primarily from Brønsted acid sites (BAS), typically present as bridging hydroxyl groups (Si–OH–Al). Under aqueous reaction conditions, these protons interact dynamically with water and alcohol molecules, leading to complex solvation and protonation behavior within confined pores. In this study, we investigate the protonation equilibrium occurring between ethanol and water at the BAS of acidic zeolites under varying hydration levels, i.e., C2H5OH–(H2O)n, n=1–4. Local structure was analyzed through an adaptive-learning global optimization algorithm, while enhanced sampling molecular dynamics simulations with Well-Tempered Metadynamics (WMetaD) and machine learning interatomic potentials (MLPs) provide free-energy surfaces (FES) at variable hydration levels. The results reveal a strong dependence of proton localization on the degree of hydration. At low hydration (1 water molecule), the proton resides predominantly on ethanol; with 2 water molecules, it shifts toward water, and at higher hydration (3 or more water molecules), it becomes extensively delocalized over the water cluster. These findings underscore the critical role of solvation in modulating acid site behavior and suggest that a minimum of three water molecules is necessary to fully stabilize the proton on water within the zeolite framework. This solvation threshold has significant implications for catalytic processes, particularly in biomass conversion reactions where alcohol protonation is a key step in dehydration mechanisms.

machine learning↗