Search NASA⌕ Search

SEARCH · Search NASA

Results for “dataset use case study”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Empirical Analysis and Automated Classification of Security Bug Reports

With the ever expanding amount of sensitive data being placed into computer systems, the need for effective cybersecurity is of utmost importance. However, there is a shortage of detailed empirical studies of security vulnerabilities from which cybersecurity metrics and best practices could be determined. This thesis has two main research goals: (1) to explore the distribution and characteristics of security vulnerabilities based on the information provided in bug tracking systems and (2) to develop data analytics approaches for automatic classification of bug reports as security or non-security related. This work is based on using three NASA datasets as case studies. The empirical analysis showed that the majority of software vulnerabilities belong only to a small number of types. Addressing these types of vulnerabilities will consequently lead to cost efficient improvement of software security. Since this analysis requires labeling of each bug report in the bug tracking system, we explored using machine learning to automate the classification of each bug report as a security or non-security related (two-class classification), as well as each security related bug report as specific security type (multiclass classification). In addition to using supervised machine learning algorithms, a novel unsupervised machine learning approach is proposed. An ac- curacy of 92%, recall of 96%, precision of 92%, probability of false alarm of 4%, F-Score of 81% and G-Score of 90% were the best results achieved during two-class classification. Furthermore, an accuracy of 80%, recall of 80%, precision of 94%, and F-score of 85% were the best results achieved during multiclass classification.

Cybersecurity↗

Assessment of ProgPy - An Open-Source Condition Monitoring and Diagnostics Tool

Traditional maintenance programs, such as corrective and preventive strategies, may lead to high costs and operational inefficiencies. Condition Monitoring and Diagnostics (CM&D) aims to improve these maintenance strategies by enabling timely insights into equipment health and performance. However, implementation of CM&D can be challenging without a robust framework that manages data efficiently, supports interoperability and simplifies integration. To address these challenges ProgPy, an open-source Python-based prognostics tool developed by NASA Ames Research Center, offers a structured solution for broader Prognostics and Health Management (PHM) applications. Ongoing research is assessing the feasibility of implementing ProgPy as a Condition Monitoring and Diagnostics (CM&D) solution by comparing its framework to the guidelines for open CM&D systems recommended in the ISO 13374-2 standard. This evaluation aims to highlight ProgPy’s strengths and identify opportunities for improvement, thereby, contributing to its advancement as an effective tool for Prognostics and Health Management (PHM). This paper presents the results of an initial assessment of the ProgPy toolbox through a gearbox case study using open-source datasets.

Condition-Monitoring, Diagnostics, Failure, Gearbo↗

A Knowledge Graph Framework for Organizing Heterogeneous Datasets for Utilization in Classical and Quantum Computing: Current Challenges and Future Directions

"The escalating impact of climate change induced extreme weather events in urban, suburban, and rural environments demands a rethink of how we have been using the single event-based or use-case-based knowledge graph models. The lack of representation in interaction within environmental variables found in literature led to the development of a novel framework that reflects the true nature of the interconnectedness in our environment. We propose an Environmental Interaction Knowledge Graph (EIKG) framework. This general EIKG framework works as the basis for interconnected environmental events by knitting interrelated events such as hurricanes leading to storm surges, which lead to flood events that could cause mudslides, landslides, etc., The cascading nature of one event leading to another related event in the environment requires an adequate understanding of each event using contextual information before conducting any data-driven analytics. This vision paper showcases how the EIKG:floods, EIKG:wildfire EIKG:landslides, etc, can be derived from a base case framework of EIKG as those individual events are interconnected with some common denominator variables. As an example, the precipitation variable is used in the flood case study as well as in the wildfire case study, as excessive precipitation levels lead to floods, and lack of precipitation leads to droughts and wildfires. We identify the precipitation variable as a “common-denominator-variable” in extreme weather events that play a key role in modeling the environment leading to different extreme weather events based on the variability of that variable (varying values where low precipitation leads to drought, and high values lead to floods). We use the insights gained from EIKG to conduct classical and Quantum Machine Learning (QML) based data analysis on the research questions developed. Our preliminary study shows how the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC) are used along with the classical machine learning models to compare the model accuracies. Our study elaborates on how a quantitative analysis uses state-of-the-art machine learning techniques that include implementing both classical and quantum machine learning models and developing the knowledge graph. The EIKG is used to organize heterogeneous datasets and integrate the relations to case-specific extreme weather events such as floods. The study uses datasets such as county-to-country residential mobility data, socioeconomic datasets from the US Census Bureau, climate and weather-related Earth Observational data from NASA, and critical infrastructure data from the Homeland Infrastructure datasets."

Knowledge Graphs, Quantum Computing, Heterogenous ↗

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

One of the largest continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available. Approaches used in Earth science research such as case study analysis and climatology studies involve gathering discovering and gathering diverse data sets and information to support the research goals. Research based on case studies involves a detailed description of specific weather events using data from different sources, to characterize physical processes in play for a specific event. Climatology-based research tends to focus on the representativeness of a given event, by studying the characteristics and distribution of a large number of events. This allows researchers to generalize characteristics such as spatio-temporal distribution, intensity, annual cycle, duration, etc. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the datasets of interest can obtain the specific files they need using these systems. However, in cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. In these cases, a search process needs to be organized around the event rather than observing instruments. In addition, the existing data systems assume users have sufficient knowledge regarding the domain vocabulary to be able to effectively utilize their catalogs. These systems do not support new or interdisciplinary researchers who may be unfamiliar with the domain terminology. This paper presents a specialized search, aggregation and curation tool for Earth science to address these existing challenges. The search tool automatically creates curated "Data Albums", aggregated collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non-relevant information and data.

Ramachandran, Rahul↗

Investigating the Response of Land-Atmosphere Interactions and Feedbacks to Spatial Representation of Irrigation in a Coupled Modeling Framework

The transport of water, heat, and momentum from the surface to the atmosphere is dependent in part on the 10 characteristics of the land surface. Together with the model physics, parameterization schemes, and parameters employed, land datasets determine the spatial variability in land surface states (i.e., soil moisture and temperature) and fluxes. Despite the importance of these datasets, they are often chosen out of convenience or regional limitations without due assessment of their impacts on model results. Irrigation is an anthropogenic form of land heterogeneity that has been shown to alter the land surface energy balance, ambient weather, and local circulations. As such, irrigation schemes are becoming more 15 prevalent in weather and climate models with rapid developments in dataset availability and parameterization scheme complexity. Thus, to address pragmatic issues related to modeling irrigation, this study uses a high-resolution, regional coupled modeling system to investigate the impacts of irrigation dataset selection on land-atmosphere (L-A) coupling using a case study from the Great Plains Irrigation Experiment (GRAINEX) field campaign. The simulations are assessed in the context of irrigated versus non-irrigated regions, subregions across the irrigation gradient, and sub-grid scale process 20 representation in coarser scale models. The results show that L-A coupling is sensitive to the choice of irrigation dataset and resolution and that the irrigation impact on surface fluxes and near surface meteorology can be dominant, conditioned on the details of the irrigation map (i.e., boundaries, heterogeneity, etc), or minimal. A consistent finding across several analyses was that even a low percentage of irrigation fraction (i.e., 4-16%) can have significant local and downstream atmospheric impacts (e.g., lower PBL height), suggesting that representation of boundaries and heterogeneous areas within irrigated 25 regions is particularly important for the modeling of irrigation impacts on the atmosphere in this model. When viewing the simulations presented here as a proxy for ‘ideal’ tiling in a Earth System Model scale gridbox, the results show that some ‘tiles’ will reach critical nonlinear moisture and planetary boundary layer (PBL) thresholds that could be important for clouds and convection, implying that heterogeneity resulting from irrigation should be taken into consideration in new sub-grid land-atmosphere exchange parameterizations.

Patricia Lawston-Parker↗

Daily evaluation of 26 precipitation datasets using Stage-IV gauge-radar data for the CONUS

New precipitation (P) datasets are released regularly, following innovations in weather forecasting models, satellite retrieval methods, and multi-source merging techniques. Using the conterminous US as a case study, we evaluated the performance of 26 gridded (sub-)daily P datasets to obtain insight into the merit of these innovations. The evaluation was performed at a daily timescale for the period 2008–2017 using the Kling–Gupta efficiency (KGE), a performance metric combining correlation, bias, and variability. As a reference, we used the high-resolution (4 km) Stage-IV gauge-radar P dataset. Among the three KGE components, the P datasets performed worst overall in terms of correlation (related to event identification). In terms of improving KGE scores for these datasets, improved P totals (affecting the bias score) and improved distribution of P intensity (affecting the variability score) are of secondary importance. Among the 11 gauge-corrected P datasets, the best overall performance was obtained by MSWEP V2.2, underscoring the importance of applying daily gauge corrections and accounting for gauge reporting times. Several uncorrected P datasets outperformed gauge-corrected ones. Among the 15 uncorrected P datasets, the best performance was obtained by the ERA5-HRES fourth-generation reanalysis, reflecting the significant advances in earth system modeling during the last decade. The (re)analyses generally performed better in winter than in summer, while the opposite was the case for the satellite-based datasets. IMERGHH V05 performed substantially better than TMPA-3B42RT V7, attributable to the many improvements implemented in the IMERG satellite P retrieval algorithm. IMERGHH V05 outperformed ERA5-HRES in regions dominated by convective storms, while the opposite was observed in regions of complex terrain. The ERA5-EDA ensemble average exhibited higher correlations than the ERA5-HRES deterministic run, highlighting the value of ensemble modeling. The WRF regional convection-permitting climate model showed considerably more accurate P totals over the mountainous west and performed best among the uncorrected datasets in terms of variability, suggesting there is merit in using high-resolution models to obtain climatological P statistics. Our findings provide some guidance to choose the most suitable P dataset for a particular application.

Hylke E. Beck↗

Alternative Datasets for Identification of Earth Science Events and Data

Alternative, or non-traditional, data sources can be used to generate datasets which can in turn be analyzed for temporal, spatial and climatological patterns. Events and case studies inferred from the analysis of these patterns can be used by the remote sensing community to more effectively search for Earth observation data. In this paper, we present a new alternative Earth science dataset created from the National Weather Service’s Area Forecast Discussion (AFD) documents. We then present an exploratory methodology for identifying interesting climatological patterns within the AFD data and a corresponding motivating example as to how these data and patterns can be used to search for relevant events or case studies.

Alternative data↗

Alternative Datasets for Identification of Earth Science Events and Data

Alternative, or non-traditional, data sources can be used to generate datasets which can in turn be analyzed for temporal, spatial and climatological patterns. Events and case studies inferred from the analysis of these patterns can be used by the remote sensing community to more effectively search for Earth observation data. In this paper, we present a new alternative Earth science dataset created from the National Weather Service’s Area Forecast Discussion (AFD) documents. We then present an exploratory methodology for identifying interesting climatological patterns within the AFD data and a corresponding motivating example as to how these data and patterns can be used to search for relevant events or case studies.

Bugbee, Kaylin↗

Data Assimilation and Reanalysis

This presentation introduces the Data Assimilation and Reanalysis principle, then details the NASA Modern-Era Retrospective analysis for Research and Applications Version 2 (MERRA-2). The MERRA-2 is atmospheric reanalysis data spanning 1980 to the present. It has been produced by the NASA Global Modeling and Assimilation Office (GMAO) and is distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). In this presentation, I will introduce the MERRA-2 datasets associated with aerosol and air quality studies and use case studies to demonstrate the data tools developed at GES DISC to analyze and visualize MERRA-2 data, such as Giovanni and Jupyter Python notebook.

Xiaohua Pan↗

Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) Datasets Released by NASA GES DISC and Their Applications for Air Quality

Nitrogen dioxide (NO2), a pervasive air pollutant, comes from vehicles, power plants, industrial emissions, and off-road sources such as construction or lawn and gardening equipment. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) curates many remote sensing datasets with NO2 retrievals, which have been utilized for air quality research and applications. The remotely-sensed datasets include those generated by the Ozone Monitoring Instrument (OMI) on the Aura satellite, the TROPOspheric Monitoring Instrument (TROPOMI) onboard the Copernicus Sentinel-5 Precursor (S5P), and the Ozone Mapping and Profiling Suite (OMPS) Nadir-Mapper (NM) instrument on the Suomi National Polar-orbiting Partnership (S- NPP). In collaboration with the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) project, the GES DISC recently released MINDS datasets. The NASA MEaSUREs MINDS project aims to develop long-term NO2 global data records by adapting a consistent retrieval algorithm to multiple instrument measurements. Long-term data records will be achieved by applying consistent retrieval approaches to multiple satellite instruments, including OMI (2004 - ); the Global Ozone Monitoring Experiment (GOME, 1995-2011) onboard the second European Remote Sensing satellite (ERS-2); the Scanning Imaging Spectrometer for Atmospheric Cartography (SCIAMACHY, 2002-2012) onboard the ENVIronmental SATellite (ENVISAT); GOME-2 on the Meteorological Operational satellites (MetOp-A and MetOp-B, 2006 - ); and TROPOMI onboard the Copernicus S5P (2017 - ). The long-term record (1995 to present) of MINDS datasets makes them very useful for air quality trend studies. Some MINDS datasets with high spatial resolution of only a few kilometers can be used for air quality research and applications at regional scales. In this presentation, we will introduce all of the MINDS products and services, and demonstrate use cases of MINDS data for studying air quality. We will also present a few other NO2 datasets acquired from NASA’s Health and Air Quality Applied Sciences Team (HAQAST), to be archived and distributed by the GES DISC, and highlight some of their applications for air quality and health.

Feng Ding↗

Urban Landscape Characterization Using Remote Sensing Data For Input into Air Quality Modeling

The urban landscape is inherently complex and this complexity is not adequately captured in air quality models that are used to assess whether urban areas are in attainment of EPA air quality standards, particularly for ground level ozone. This inadequacy of air quality models to sufficiently respond to the heterogeneous nature of the urban landscape can impact how well these models predict ozone pollutant levels over metropolitan areas and ultimately, whether cities exceed EPA ozone air quality standards. We are exploring the utility of high-resolution remote sensing data and urban growth projections as improved inputs to meteorological and air quality models focusing on the Atlanta, Georgia metropolitan area as a case study. The National Land Cover Dataset at 30m resolution is being used as the land use/land cover input and aggregated to the 4km scale for the MM5 mesoscale meteorological model and the Community Multiscale Air Quality (CMAQ) modeling schemes. Use of these data have been found to better characterize low density/suburban development as compared with USGS 1 km land use/land cover data that have traditionally been used in modeling. Air quality prediction for future scenarios to 2030 is being facilitated by land use projections using a spatial growth model. Land use projections were developed using the 2030 Regional Transportation Plan developed by the Atlanta Regional Commission. This allows the State Environmental Protection agency to evaluate how these transportation plans will affect future air quality.

Quattrochi, Dale A.↗

The NASA Merra-2 Reanalysis Products: Data and Tools Used for Aerosol and Air Quality Studies

The NASA Modern-Era Retrospective analysis for Research and Applications Version 2 (MERRA-2) is atmospheric reanalysis data spanning 1980 to present. It has been produced by the NASA Global Modeling and Assimilation Office (GMAO) and is distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). MERRA-2 data includes 100 collections of Earth system variables, mainly from the atmospheric model, such as aerosol fields and meteorological fields, radiation fields, and aerosol fields, guided by the assimilation of as many as six million observations every six hours. MERRA-2 has been one of the most popular datasets from NASA and is widely used in interdisciplinary research and applications, with increasing numbers of new users. For example, at least 7000 users accessed MERRA-2 data at GES DISC in the year 2021, ~1000 more users than in the year 2020. In this presentation, we will introduce the MERRA-2 datasets associated with aerosol and air quality studies and use a wildfire case study to demonstrate the data tools developed at GES DISC to analyze and visualize MERRA-2 data, such as Giovanni and the level 3 and level 4 subsetter, and Jupyter Python notebook. We will also update the status of cloud migration of the MERRA-2 data to Amazon Web Services (AWS).

Xiaohua Pan↗

Evaluating the Impacts of NASA/SPoRT Daily Greenness Vegetation Fraction on Land Surface Model and Numerical Weather Forecasts

The NASA Short-term Prediction Research and Transition (SPoRT) Center has developed a Greenness Vegetation Fraction (GVF) dataset, which is updated daily using swaths of Normalized Difference Vegetation Index data from the Moderate Resolution Imaging Spectroradiometer (MODIS) data aboard the NASA EOS Aqua and Terra satellites. NASA SPoRT began generating daily real-time GVF composites at 1-km resolution over the Continental United States (CONUS) on 1 June 2010. The purpose of this study is to compare the National Centers for Environmental Prediction (NCEP) climatology GVF product (currently used in operational weather models) to the SPoRT-MODIS GVF during June to October 2010. The NASA Land Information System (LIS) was employed to study the impacts of the SPoRT-MODIS GVF dataset on a land surface model (LSM) apart from a full numerical weather prediction (NWP) model. For the 2010 warm season, the SPoRT GVF in the western portion of the CONUS was generally higher than the NCEP climatology. The eastern CONUS GVF had variations both above and below the climatology during the period of study. These variations in GVF led to direct impacts on the rates of heating and evaporation from the land surface. In the West, higher latent heat fluxes prevailed, which enhanced the rates of evapotranspiration and soil moisture depletion in the LSM. By late Summer and Autumn, both the average sensible and latent heat fluxes increased in the West as a result of the more rapid soil drying and higher coverage of GVF. The impacts of the SPoRT GVF dataset on NWP was also examined for a single severe weather case study using the Weather Research and Forecasting (WRF) model. Two separate coupled LIS/WRF model simulations were made for the 17 July 2010 severe weather event in the Upper Midwest using the NCEP and SPoRT GVFs, with all other model parameters remaining the same. Based on the sensitivity results, regions with higher GVF in the SPoRT model runs had higher evapotranspiration and lower direct surface heating, which typically resulted in lower (higher) predicted 2-m temperatures (2-m dewpoint temperatures). Portions of the Northern Plains states experienced substantial increases in convective available potential energy as a result of the higher SPoRT/MODIS GVFs. These differences produced subtle yet quantifiable differences in the simulated convective precipitation systems for this event.

Bell, Jordan R.↗

VESIcal: A Critical Approach to Volatile Solubility Modelling Using the Open-Source Engine Vesical

Accurate models of H(2)O and CO(2) solubility in silicate melts are vital for understanding volcanic plumbing systems. These models are used to estimate the depths of magma storage regions from melt inclusion volatile contents, investigate the role of volatile exsolution as a driver of volcanic eruptions, and track the degassing path followed by a magma ascending to the surface. However, despite the large increase in the number of experimental constraints over the last two decades, many recent studies still utilize an earlier generation of models which were calibrated on experimental datasets with restricted compositional ranges. This may be because many of the available tools for more recent models require large numbers of input parameters to be hand-typed (e.g., temperature, concentrations of H(2)O, CO(2), and 8–14 oxides), making them difficult to implement on large datasets. Here, we use a new open-source Python3 tool, VESIcal, to critically evaluate the behaviors and sensitivities of different solubility models for a range of melt compositions. Using literature datasets of andesitic-dacitic experimental products and melt inclusions as case studies, we illustrate the importance of evaluating the calibration dataset of each model. Finally, we highlight the limitations of particular data presentation methods, such as isobar diagrams, and provide suggestions for alternatives, and best practices regarding the presentation and archiving of data. This review will aid the selection of the most applicable solubility model for different melt compositions, and identifies areas where additional experimental constraints on volatile solubility are required.

magma↗

Drought Monitoring for 3 North American Case Studies Based on the North American Land Data Assimilation System (NLDAS)

Both NLDAS Phase 1 (1996-2007) and Phase 2 (1979-present) datasets have been evaluated against in situ observational datasets, and NLDAS forcings and outputs are used by a wide variety of users. Drought indices and drought monitoring from NLDAS were recently examined by Mo et al. (2010) and Sheffield et al. (2010). In this poster, we will present results analyzing NLDAS Phase 2 forcings and outputs for 3 North American Case studies being analyzed as part of the NOAA MAPP Drought Task Force: (1) Western US drought (1998- 2004); (2) plains/southeast US drought (2006-2007); and (3) Current Texas-Mexico drought (2011-). We will examine percentiles of soil moisture consistent with the NLDAS drought monitor.

Peters-Lidard, Christa D.↗

An Integrated Examination of AMSR2 Products over Ocean

Integrated processing and analysis are presented for oceanic retrievals of precipitation, sea ice concentration, columnar water vapor and cloud water, sea surface temperature, and near surface wind from the advanced microwave scanning radiometer-2 (AMSR2) sensor. By developing a common algorithmic framework and permitting iterative interaction between historically separate science algorithms, ambiguous and contradictory retrieval results are minimized. The integration also serves to improve each algorithm individually, decreasing reliance on ancillary datasets. Case studies are presented that exemplify both ongoing challenges for these retrievals and potential research uses of such an integrated satellite product. Two cases are presented each for precipitation at high latitudes and marginal sea ice detection, with analysis supplemented by satellite data from CloudSat profiles and Himawari-8 imagery. Detection and retrieval of snowfall remains a challenge, while light rainfall detection can be aided by a variational algorithm. Sea ice concentrations below 20% cause disagreements, but the algorithms otherwise agree well on the sea ice edge. Potential mitigation strategies for ambiguous areas of light rainfall and marginal sea ice are discussed. The analysis demonstrates potential avenues for future algorithm development, but also some physical limitations of remote sensing with the AMSR2 frequencies.

variational methods↗

Assessment of NASA's Physiographic and Meteorological Datasets as Input to HSPF and SWAT Hydrological Models

This paper documents the use of simulated Moderate Resolution Imaging Spectroradiometer land use/land cover (MODIS-LULC), NASA-LIS generated precipitation and evapo-transpiration (ET), and Shuttle Radar Topography Mission (SRTM) datasets (in conjunction with standard land use, topographical and meteorological datasets) as input to hydrological models routinely used by the watershed hydrology modeling community. The study is focused in coastal watersheds in the Mississippi Gulf Coast although one of the test cases focuses in an inland watershed located in northeastern State of Mississippi, USA. The decision support tools (DSTs) into which the NASA datasets were assimilated were the Soil Water & Assessment Tool (SWAT) and the Hydrological Simulation Program FORTRAN (HSPF). These DSTs are endorsed by several US government agencies (EPA, FEMA, USGS) for water resources management strategies. These models use physiographic and meteorological data extensively. Precipitation gages and USGS gage stations in the region were used to calibrate several HSPF and SWAT model applications. Land use and topographical datasets were swapped to assess model output sensitivities. NASA-LIS meteorological data were introduced in the calibrated model applications for simulation of watershed hydrology for a time period in which no weather data were available (1997-2006). The performance of the NASA datasets in the context of hydrological modeling was assessed through comparison of measured and model-simulated hydrographs. Overall, NASA datasets were as useful as standard land use, topographical , and meteorological datasets. Moreover, NASA datasets were used for performing analyses that the standard datasets could not made possible, e.g., introduction of land use dynamics into hydrological simulations

Alacron, Vladimir J.↗