Search NASA⌕ Search

SEARCH · Search NASA

Results for “exploratory data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

Experience of the JPL Exploratory Data Analysis Team at validating HIRS2/MSU cloud parameters

Validation of the HIRS2/MSU cloud parameters began with the cloud/climate feedback problem. The derived effective cloud amount is less sensitive to surface temperature for higher clouds. This occurs because as the cloud elevation increases, the difference between surface temperature and cloud temperature increases, so only a small change in cloud amount is needed to effect a large change in radiance at the detector. By validating the cloud parameters it is meant 'developing a quantitative sense for the physical meaning of the measured parameters', by: (1) identifying the assumptions involved in deriving parameters from the measured radiances, (2) testing the input data and derived parameters for statistical error, sensitivity, and internal consistency, and (3) comparing with similar parameters obtained from other sources using other techniques.

Kahn, Ralph↗

DataHub--Knowledge-Based Science Data Management for Exploratory Data Analysis

It is our belief that new modes of research and new tools will be required to handle the massive amount of diverse data that is to be stored, organized, accessed, distributed, visualized, and analyzed. The fundamental innovation required is the integration of three automation technologies, videlicet knowledge-based expert systems, science visualization and science data management. This integration is based on a concept caled the DataHub, which we describe here.

DataHub↗

The crush of new earth science data knocking at our door

The reasons for collecting massive amounts of earth science data in the Earth Observing System (EOS) Project are discussed. A processing hierarchy for handling the data is described, and the prospects for adequate throughput and storage for operational data analysis in the EOS era are addressed. Needs for successful exploratory data analysis are examined, and the policy issues implicated by the large stream of EOS data are considered.

Kahn, Ralph↗

Machine learning to discover mineral trapping signatures due to CO 2 injection

Mineral trapping is pursued as a geological CO 2 sequestration (GCS) mechanism because it permanently stores CO 2 in solid phases or minerals. However, CO 2 mineral-trapping mechanisms are poorly understood due to (1) lack of sufficient field and laboratory data characterizing these complex processes, and (2) challenges to develop site-specific reactive-transport models coupling fluid flow and geochemical reactions occurring at various temporal (from milliseconds to years) and spatial (from pore (millimeters) to field (kilometers)) scales. Reactive transport with additional complexities such as heterogeneity can make the simulation outputs even more difficult to interpret because of complex nonlinearity and multi-scale interdependencies. Furthermore, the values of model outputs such as concentrations can vary by several orders of magnitude, making it harder to correlate and characterize the impact of the variables via traditional data interpretation techniques such as exploratory data analyses. Recently, machine learning (ML) has shown promise in feature discovery and in highlighting hidden mechanisms that cannot be obtained by existing data-analytics and statistical methods. In this study, we applied an unsupervised ML approach, non-negative matrix factorization with custom -means clustering (NMF) to the data generated by reactive-transport simulations of GCS. The reactive-transport data consisted of 19 attributes, including four physio-chemical variables (pH, porosity, aqueous CO 2 , and sequestered CO 2 ), six chemical species (K + , Na + , HCO, Ca 2+ , Mg 2+ , Fe 2+ ), and four carbonate minerals (calcite, dolomite, siderite, and ankerite), a feldspar mineral (albite), and four clay minerals (illite, clinochlore, kaolinite, and smectite) over a period of 200 years of simulation time. Furthermore, the simulation data used was for Morrow B sandstone at the Farnsworth hydrocarbon unit in Texas. Data are sampled at two locations within the model domain: (1) at the injection well and (2) 200 m west of the injection well. The injection was performed for a period of 10 years. Using NMF, we estimated the temporal interdependencies among the 19 attributes over a span of 200 years. We found that NMF was able to identify four reaction stages and their dominant attributes; these cannot be directly discerned through traditional visualization (e.g., line plots, Pareto analysis, Glyph-based visualization methods) or exploratory data analysis tools of the simulation data. The four stages were: reactions in the injection phase followed by short-, mid-, and long-term reactions. The NMF analysis also revealed that 10 among the 19 attributes are dominant. These dominant attributes for mineral trapping include calcite, dolomite at injection well, siderite at 200 m away from the injection well, clinochlore, kaolinite, Na + , K + , Ca 2+ , Mg 2+ , pH, and aqeuous CO 2 . Finally, at late times (65–200 years), our results showed that calcite plays a major role in mineral trapping with insignificant contribution from siderite, ankerite, and clay minerals. These findings make the proposed unsupervised ML-model attractive for reactive-transport sensing towards real-time GCS monitoring.

54 ENVIRONMENTAL SCIENCES↗

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION↗

Exploratory Climate Data Visualization and Analysis Using DV3D and UVCDAT

Earth system scientists are being inundated by an explosion of data generated by ever-increasing resolution in both global models and remote sensors. Advanced tools for accessing, analyzing, and visualizing very large and complex climate data are required to maintain rapid progress in Earth system research. To meet this need, NASA, in collaboration with the Ultra-scale Visualization Climate Data Analysis Tools (UVCOAT) consortium, is developing exploratory climate data analysis and visualization tools which provide data analysis capabilities for the Earth System Grid (ESG). This paper describes DV3D, a UV-COAT package that enables exploratory analysis of climate simulation and observation datasets. OV3D provides user-friendly interfaces for visualization and analysis of climate data at a level appropriate for scientists. It features workflow inte rfaces, interactive 40 data exploration, hyperwall and stereo visualization, automated provenance generation, and parallel task execution. DV30's integration with CDAT's climate data management system (COMS) and other climate data analysis tools provides a wide range of high performance climate data analysis operations. DV3D expands the scientists' toolbox by incorporating a suite of rich new exploratory visualization and analysis methods for addressing the complexity of climate datasets.

Maxwell, Thomas↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Uncovering interpretable relationships in high-dimensional scientific data through function preserving projections

Abstract In many fields of science and engineering, we frequently encounter experiments or simulations datasets that describe the behavior of complex systems and uncovering human interpretable patterns between their inputs and outputs via exploratory data analysis is essential for building intuition and facilitating discovery. Often, we resort to 2D embeddings for examining these high-dimensional relationships (e.g. dimensionality reduction). However, most existing embedding methods treat the dimensions as coordinates for samples in a high-dimensional space, which fail to capture the potential functional relationships, and the few methods that do take function into consideration either only focus on linear patterns or produce non-linear embeddings that are hard to interpret. To address these challenges, we proposed function preserving projections (FPP), which construct 2D linear embeddings optimized to reveal interpretable yet potentially non-linear patterns between the domain and the range of a high-dimensional function. The intuition here is that humans are good at understanding potentially non-linear patterns in 2D but unable to interpret non-linear mapping from high-dimensional space to 2D. Therefore, we should restrict the projection to linear but not the pattern we are seeking. Using FPP on real-world datasets, one can obtain fundamentally new insights about high-dimensional relationships in extremely large datasets that could not be processed with existing dimension reduction methods.

97 MATHEMATICS AND COMPUTING↗

Automated analysis of complex data

We have examined some of the issues involved in automating exploratory data analysis, in particular the tradeoff between control and opportunism. We have proposed an opportunistic planning solution for this tradeoff, and we have implemented a prototype, Igor, to test the approach. Our experience in developing Igor was surprisingly smooth. In contrast to earlier versions that relied on rule representation, it was straightforward to increment Igor's knowledge base without causing the search space to explode. The planning representation appears to be both general and powerful, with high level strategic knowledge provided by goals and plans, and the hooks for domain-specific knowledge are provided by monitors and focusing heuristics.

Saintamant, Robert↗

Knowledge-based assistance for science visualization and analysis using large distributed databases

Within this decade, the growth in complexity of exploratory data analysis and the sheer volume of space data require new and innovative approaches to support science investigators in achieving their research objectives. To date, there have been numerous efforts addressing the individual issues involved in inter-disciplinary, multi-instrument investigations. However, while successful in small scale, these efforts have not proven to be open and scaleable. This proposal addresses four areas of significant need: scientific visualization and analysis; science data management; interactions in a distributed, heterogeneous environment; and knowledge-based assistance for these functions. The fundamental innovation embedded within this proposal is the integration of three automation technologies, namely, knowledge-based expert systems, science visualization and science data management. This integration is based on the concept called the Data Hub. With the Data Hub concept, NASA will be able to apply a more complete solution to all nodes of a distributed system. Both computation nodes and interactive nodes will be able to effectively and efficiently use the data services (access, retrieval, update, etc.) with a distributed, interdisciplinary information system in a uniform and standard way. This will allow the science investigators to concentrate on their scientific endeavors, rather than to involve themselves in the intricate technical details of the systems and tools required to accomplish their work. Thus, science investigators need not be programmers. The emphasis will be on the definition and prototyping of system elements with sufficient detail to enable data analysis and interpretation leading to publishable scientific results. In addition, the proposed work includes all the required end-to-end components and interfaces to demonstrate the completed concept.

Handley, Thomas H., Jr.↗

Knowledge-based assistance for science visualization and analysis using large distributed databases

Within this decade, the growth in complexity of exploratory data analysis and the sheer volume of space data require new and innovative approaches to support science investigators in achieving their research objectives. To date, there have been numerous efforts addressing the individual issues involved in inter-disciplinary, multi-instrument investigations. However, while successful in small scale, these efforts have not proven to be open and scalable. This proposal addresses four areas of significant need: scientific visualization and analysis; science data management; interactions in a distributed, heterogeneous environment; and knowledge-based assistance for these functions. The fundamental innovation embedded with this proposal is the integration of three automation technologies, namely, knowledge-based expert systems, science visualization and science data management. This integration is based on concept called the DataHub. With the DataHub concept, NASA will be able to apply a more complete solution to all nodes of a distributed system. Both computation nodes and interactive nodes will be able to effectively and efficiently use the data services (address, retrieval, update, etc.) with a distributed, interdisciplinary information system in a uniform and standard way. This will allow the science investigators to concentrate on their scientific endeavors, rather than to involve themselves in the intricate technical details of the systems and tools required to accomplish their work. Thus, science investigators need not be programmers. The emphasis will be on the definition and prototyping of system elements with sufficient detail to enable data analysis and interpretation leading to publishable scientific results. In addition, the proposed work includes all the required end-to-end components and interfaces to demonstrate the completed concept.

Handley, Thomas H., Jr.↗

Radioisotope Identification with List-Mode Gamma-Ray Data

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with spectra containing similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and qualitative distribution analysis. Additionally, we propose a probabilistic classification model that can utilize spectral data, temporal data, or both to determine if the incorporation of temporal information improves radioisotope identification. Our findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated to develop more robust and optimal methods for utilizing this temporal information in applications requiring radioisotope identification.

List-mode data↗

Exploring Explicit Uncertainty for Binary Analysis (EUBA)

Reverse engineering (RE) analysts struggle to address critical questions about the safety of binary code accurately and promptly, and their supporting program analysis tools are simply wrong sometimes. The analysis tools have to approximate in order to provide any information at all, but this means that they introduce uncertainty into their results. And those uncertainties chain from analysis to analysis. We hypothesize that exposing sources, impacts, and control of uncertainty to human binary analysts will allow the analysts to approach their hardest problems with high-powered analytic techniques that they know when to trust. Combining expertise in binary analysis algorithms, human cognition, uncertainty quantification, verification and validation, and visualization, we pursue research that should benefit binary software analysis efforts across the board. We find a strong analogy between RE and exploratory data analysis (EDA); we begin to characterize sources and types of uncertainty found in practice in RE (both in the process and in supporting analyses); we explore a domain-specific focus on uncertainty in pointer analysis, showing that more precise models do help analysts answer small information flow questions faster and more accurately; and we test a general population with domain-general sudoku problems, showing that adding "knobs" to an analysis does not significantly slow down performance. This document describes our explorations in uncertainty in binary analysis.

97 MATHEMATICS AND COMPUTING↗

Radioisotope Identification with List-Mode Gamma Ray Data: A rigorous assessment on the value of temporal information applied to radioisotope identification.

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with “confuser” sources, or spectra with similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research rigorously examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and information theory. We further propose a basic classification model that can utilize spectral or temporal data (or both) to determine if the incorporation of temporal information can improve radioisotope identification. The findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗