Search NASASearch

Engineering topics

Jennifer C Wei

Publications and source records attributed to Jennifer C Wei.

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source

PBL Height from AIRS, GPS RO, and MERRA-2 Products in NASA GES DISC and Their 10 Year Seasonal Mean Intercomparison

Within the planetary boundary layer (PBL), surface forcing response, drag, turbulence, and vertical mixing are important processes and play a more critical role here than in the overlying “free atmosphere”. The PBL Height (PBLH) is an important parameter in climate models, weather forecasts, and air quality prediction. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) provides data processing, archiving, and distribution services for numerous Earth science products. PBLH is a parameter in three products served by GES DISC, which are from the Atmospheric Infrared Sounder (AIRS), the Global Positioning System (GPS) radio occultation (RO) experiment, and the NASA reanalysis product Modern-Era Retrospective analysis for Research and Applications – 2 (MERRA-2). These products have different spatial and temporal resolutions and coverages, and their PBLH definitions are also different. To better serve the PBL research community, we have summarized the specifications of these products. A ten-year seasonal mean intercomparison is also conducted to provide further guidance to users. The intercomparison results show that MERRA-2 has a much shallower PBL than AIRS and GPS RO. An experimental study indicates the different PBLH definition in MERRA-2 caused smaller values of PBLH. The improvement of the water vapor retrieval in AIRS version 7 over version 6 results in the version 7 PBLH agreeing better with GPS RO and MERRA-2 than version 6, especially near the equator and low latitudes.

Feng Ding

Exploring Anomalous PM 2.5 from Wildfires and Dust Storms using Data and Services at NASA GES DISC

The presence of fine particles in the atmosphere with a diameter of less than 2.5 µm, called particulate matter 2.5 (PM 2.5 ), poses a significant threat to human health as a criteria air pollutant. Fortunately, NASA's Goddard Earth Sciences Data and Information Services Center (GES DISC) provides easy access to several PM 2.5 concentration products. These datasets include the reanalysis of global hourly and monthly aerosol components including PM 2.5 data from the Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA-2), as well as 3-hourly real-time ensemble forecasts of PM 2.5 from the Hazardous Air Quality Ensemble System (HAQES). The HAQES products are developed by the George Mason University Air Quality Laboratory as part of NASA's Health Air Quality Applied Science Team (HAQAST). The GES DISC is actively collaborating with scientists in the HAQAST program to further expand air quality data collections. Two new datasets are currently being archived: one is the machine learning-based global hourly PM 2.5 derived from MERRA-2; the other is the localized data (NO 2 , O 3 , and PM 2.5 ) time series derived from NASA's GEOS Composition Forecasting (GEOS-CF) system. In this presentation, we will explore the spatial patterns and long-distance transport characteristics of elevated PM 2.5 during extreme pollution events, such as the June 2023 Canadian wildfires, which are still active at the time of writing; and severe spring dust storms in 2023 over Asia. To gain comprehensive insights, we will utilize various PM 2.5 data in conjunction with satellite-observed aerosol data from TROPOspheric Monitoring Instrument (TROPOMI) on Sentinel-5P. The primary focus of this presentation will be to demonstrate effective use of data tools and services to visualize and explore extreme air pollution phenomena. Additionally, we will provide guidance on how users can download specific data of interest, facilitating further analysis and research in this critical area.

air quality

Quantum Leap: Evaluating the Feasibility of Quantum Machine Learning Using NASA Earth Observational Data

This study explores the feasibility of leveraging quantum machine learning (QML) to analyze NASA Earth Observational (EO) data for climate change research, with a particular focus on the phenomenon of ”crop frosting” which has become more prevalent due to climate change. We implemented and evaluated two QML models, the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC), in both simulated and real quantum computing environments using a 127 qubit IBM quantum processor. Our study emphasizes the scientific rigor in comparing these quantum models with a classical Support Vector Machine (SVM) classifier, highlighting their performance in processing climate data. The results offer valuable insights into the potential scientific advantages, limitations, and scalability of QML for analyzing EO datasets, thus paving the way for more advanced climate modeling and predictive analytics using quantum computing. We showcased how Environmental Interaction Knowledge Graphs (EIKGs) and Digital Twins (DTs) can be integrated into this study. This research underscores the transformative potential of Classical and QML leveraging KGs and DT to address the multifaceted challenges posed by climate change.

Quantum Computing

Natural Language Processing for Extracting Rich Disease Data Aligned To Satellite Meteorological Data

Global climate change is redefining our understanding of how diseases spread. In Sri Lanka, vector-borne diseases such as dengue fever, encephalitis, and leptospirosis historically surged during the monsoon seasons when temperatures were high enough for mosquito eggs to hatch. Unfortunately, due to rising temperatures and more erratic rainfall patterns, mosquito eggs can now hatch year-round and are increasingly unpredictable, leading to an alarmingly increasing number of hospitalizations and deaths. More data is needed to adapt our response to these diseases in an increasingly warmer world. In the contemporary landscape, a wealth of disease information is available, yet accessibility remains limited due to unstructured data formats such as PDFs. Therefore, converting unstructured disease reports into structured formats is necessary for effectively leveraging data. This paper introduces a comprehensive framework for collecting unstructured disease reports and transforming them into analyzable formats. By creating separate models tailored to each data format, we can ensure accuracy compared to general models. These straightforward models enhance accessibility and empower other researchers to use our tools. The returned structured data can then be harnessed for analysis, statistical purposes, and informing evidence-based public health interventions, thus facilitating more informed decision-making in healthcare. We deploy this framework to produce geospatial data for Sri Lanka and Brazil for many different conditions and align these data with satellite environmental data, providing for the first time a structured, aligned powerful dataset for disease modeling.

Open Source Open Science

Graph Representation Learning for Dengue Forecasting

In 2017, the largest recorded dengue outbreak in Sri Lanka’s history occurred. Since then, dengue has continued to threaten national health across Sri Lanka. The development of an effective Early Warning System (EWS) for dengue outbreaks is essential for Sri Lanka’s Ministry of Health to take preventative measures. We propose the use of Graph Neural Networks as EWS. Using earth observational data from NASAs global satellites and dengue incidence data from Sri Lanka s Ministry of Health, we developed a series of traditional and graph representation EWS to forecast Dengue cases across Sri Lanka’s 25 districts between 2013 and 2022. We demonstrate empirically that Graph Neural Networks which incorporate spatiotemporal relations significantly outperform traditional EWS such as Autoregressive Integrated Moving Average (ARIMA), Random Forest, and Long Short-Term Memory (LSTM). Our source code is available on GitHub and will be provided in the final submission.

Graph Neural Networks

A Knowledge Graph Framework for Organizing Heterogeneous Datasets for Utilization in Classical and Quantum Computing: Current Challenges and Future Directions

"The escalating impact of climate change induced extreme weather events in urban, suburban, and rural environments demands a rethink of how we have been using the single event-based or use-case-based knowledge graph models. The lack of representation in interaction within environmental variables found in literature led to the development of a novel framework that reflects the true nature of the interconnectedness in our environment. We propose an Environmental Interaction Knowledge Graph (EIKG) framework. This general EIKG framework works as the basis for interconnected environmental events by knitting interrelated events such as hurricanes leading to storm surges, which lead to flood events that could cause mudslides, landslides, etc., The cascading nature of one event leading to another related event in the environment requires an adequate understanding of each event using contextual information before conducting any data-driven analytics. This vision paper showcases how the EIKG:floods, EIKG:wildfire EIKG:landslides, etc, can be derived from a base case framework of EIKG as those individual events are interconnected with some common denominator variables. As an example, the precipitation variable is used in the flood case study as well as in the wildfire case study, as excessive precipitation levels lead to floods, and lack of precipitation leads to droughts and wildfires. We identify the precipitation variable as a “common-denominator-variable” in extreme weather events that play a key role in modeling the environment leading to different extreme weather events based on the variability of that variable (varying values where low precipitation leads to drought, and high values lead to floods). We use the insights gained from EIKG to conduct classical and Quantum Machine Learning (QML) based data analysis on the research questions developed. Our preliminary study shows how the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC) are used along with the classical machine learning models to compare the model accuracies. Our study elaborates on how a quantitative analysis uses state-of-the-art machine learning techniques that include implementing both classical and quantum machine learning models and developing the knowledge graph. The EIKG is used to organize heterogeneous datasets and integrate the relations to case-specific extreme weather events such as floods. The study uses datasets such as county-to-country residential mobility data, socioeconomic datasets from the US Census Bureau, climate and weather-related Earth Observational data from NASA, and critical infrastructure data from the Homeland Infrastructure datasets."

Knowledge Graphs, Quantum Computing, Heterogenous

Exploring Flooded Fraction Prediction through Machine Learning Models Focusing on Medical Infrastructure in the Southeast U.S. Coastal Areas

Rising sea levels due to climate change increasingly threaten medical infrastructure through flooding. This study develops machine learning models to predict flood exposure for 11,508 medical facilities in the southeastern coastal regions of the United States by integrating datasets including meteorological, hydrological, topographic, and geological data, the Natural Risk Index, and historical flood records from NASA, HIFLD, and FEMA. Six regression models, namely Linear Regression, Support Vector Regression, Random Forest, k-Nearest Neighbors, XGBoost, and Artificial Neural Networks, are trained using 16 explanatory variables identified through literature review and correlation analysis. Data preprocessing employs the SMOGN for class imbalance and Winsorization for outliers. Model performance is evaluated using MAE, MSE, and RMSE, with Random Forest and XGBoost models achieving the highest performance (MSE of 2.58e-5 and 3.69e-5, respectively). This multifactorial approach allows the models to capture complex flood-influencing relationships, enhancing adaptability and performance across geographic regions. Future work focuses on expanding across the U.S. and developing a near real-time flood monitoring system.

Jihoon Chung