Search NASASearch

SEARCH · Search NASA

Results for “high dimensional data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Topics in inference and decision-making with partial knowledge

Two essential elements needed in the process of inference and decision-making are prior probabilities and likelihood functions. When both of these components are known accurately and precisely, the Bayesian approach provides a consistent and coherent solution to the problems of inference and decision-making. In many situations, however, either one or both of the above components may not be known, or at least may not be known precisely. This problem of partial knowledge about prior probabilities and likelihood functions is addressed. There are at least two ways to cope with this lack of precise knowledge: robust methods, and interval-valued methods. First, ways of modeling imprecision and indeterminacies in prior probabilities and likelihood functions are examined; then how imprecision in the above components carries over to the posterior probabilities is examined. Finally, the problem of decision making with imprecise posterior probabilities and the consequences of such actions are addressed. Application areas where the above problems may occur are in statistical pattern recognition problems, for example, the problem of classification of high-dimensional multispectral remote sensing image data.

Safavian, S. Rasoul

Analyzing Tropical Waves Using the Parallel Ensemble Empirical Model Decomposition Method: Preliminary Results from Hurricane Sandy

In this study, we discuss the performance of the parallel ensemble empirical mode decomposition (EMD) in the analysis of tropical waves that are associated with tropical cyclone (TC) formation. To efficiently analyze high-resolution, global, multiple-dimensional data sets, we first implement multilevel parallelism into the ensemble EMD (EEMD) and obtain a parallel speedup of 720 using 200 eight-core processors. We then apply the parallel EEMD (PEEMD) to extract the intrinsic mode functions (IMFs) from preselected data sets that represent (1) idealized tropical waves and (2) large-scale environmental flows associated with Hurricane Sandy (2012). Results indicate that the PEEMD is efficient and effective in revealing the major wave characteristics of the data, such as wavelengths and periods, by sifting out the dominant (wave) components. This approach has a potential for hurricane climate study by examining the statistical relationship between tropical waves and TC formation.

PEEMD

Analysis of Ice Mass Growth Over Time on the CRM65 Midspan Hybrid Model

The Aeronautics Research Mission Directorate at NASA is developing and applying tools to enable future technologies towards sustainable flight. Aircraft icing has been identified as a potential barrier to entry into service for innovative designs necessitating improvements to computational ice accretion tools. NASA is developing the Glenn Icing Computational Environment (GlennICE) to address deficiencies in the computational modeling capabilities of previously developed ice accretion solvers. To benchmark and improve the ability to model highly three-dimensional ice accretion, high quality validation data against experimental data is required. The CRM65 Midspan Hybrid geometry was previously tested at the NASA Icing Research Tunnel to generate experimental data for swept wing geometries typical for commercial transport aircraft. As a part of a 2018 icing test campaign, experimental data characterizing the relationship between ice accretion time and ice mass growth was obtained and can be leveraged for use in validation of computational tools. The desire for computational ice accretion solvers to predict ice shapes profiles accreted experimentally has often overshadowed the comparison to the mass and bulk volume of ice accreted. To address this deficiency, an analysis is presented in which GlennICE is applied to simulations of the CRM65 Midspan Hybrid model tested in the NASA Icing Research Tunnel. Results from the computational fluid dynamics simulations compared favorably to the experimental pressure coefficient data, thus validating the modeling setup. The experimental data showed excellent repeatability for the 15.0 minute accretion time. The comparisons between the experimental and computational ice mass over time showed good agreement up to 10.0 minutes after which the ice mass was underpredicted. The experimental ice mass was largely linear with some nonlinear data. The bulk volume of ice accreted experimentally compared well to GlennICE for the scanned ice shapes and mean combined cross section ice shapes, but was underpredicted for the maximum combined cross section ice shapes at longer accretion times. The experimental minimum combined cross section, mean combined cross section, and maximum combined cross section profiles when compared to GlennICE show good agreement for the mean combined cross section up to 15.0 minutes. The analyses show that with a single-shot method, GlennICE currently underpredicts the ice mass for longer accretion times, is not able to match the bulk volume of the maximum combined cross section due to dominating scallop features, and future work is required to generate a more generalized ice bulk density model.

Icing

Analysis of Ice Mass Growth Over Time on the CRM65 Midspan Hybrid Model

The Aeronautics Research Mission Directorate at NASA is developing and applying tools to enable future technologies towards sustainable flight. Aircraft icing has been identified as a potential barrier to entry into service for innovative designs necessitating improvements to computational ice accretion tools. NASA is developing the Glenn Icing Computational Environment (GlennICE) to address deficiencies in the computational modeling capabilities of previously developed ice accretion solvers. To benchmark and improve the ability to model highly three-dimensional ice accretion, high quality validation data against experimental data is required. The CRM65 Midspan Hybrid geometry was previously tested at the NASA Icing Research Tunnel to generate experimental data for swept wing geometries typical for commercial transport aircraft. As a part of a 2018 icing test campaign, experimental data characterizing the relationship between ice accretion time and ice mass growth was obtained and can be leveraged for use in validation of computational tools. The desire for computational ice accretion solvers to predict ice shapes profiles accreted experimentally has often overshadowed the comparison to the mass and bulk volume of ice accreted. To address this deficiency, an analysis is presented in which GlennICE is applied to simulations of the CRM65 Midspan Hybrid model tested in the NASA Icing Research Tunnel. Results from the computational fluid dynamics simulations compared favorably to the experimental pressure coefficient data, thus validating the modeling setup. The experimental data showed excellent repeatability for the 15.0 minute accretion time. The comparisons between the experimental and computational ice mass over time showed good agreement up to 10.0 minutes after which the ice mass was underpredicted. The experimental ice mass was largely linear with some nonlinear data. The bulk volume of ice accreted experimentally compared well to GlennICE for the scanned ice shapes and mean combined cross section ice shapes, but was underpredicted for the maximum combined cross section ice shapes at longer accretion times. The experimental minimum combined cross section, mean combined cross section, and maximum combined cross section profiles when compared to GlennICE show good agreement for the mean combined cross section up to 15.0 minutes. The analyses show that with a single-shot method, GlennICE currently underpredicts the ice mass for longer accretion times, is not able to match the bulk volume of the maximum combined cross section due to dominating scallop features, and future work is required to generate a more generalized ice bulk density model.

Icing

Affine Transformations to Enable Machine Learning for Semi-Quantitative EDS Analysis

Energy Dispersive X-ray Spectroscopy (EDS) is an essential technique for determining elemental concentrations and distributions within microstructures, critical for materials discovery, optimization, and qualification. However, most published EDS data is qualitative because current quantitative EDS analysis methods require extensive calibration and post-processing, limiting their practicality and widespread adoption. This work seeks to establish a framework for accelerated EDS characterization and spectrum analysis that can leverage ML to analyze correlations between various elemental compositions and resulting EDS spectra. The complex physics and data result in a high-dimensional problem that grows exponentially with the number of elements in the system and the complexity of the spectrum analysis. ML provides a way to compute and optimize the results of this highly dimensional problem in a flexible way to tailor it to the user’s specific needs and material system. However, the framework emphasizes transparency through a strictly mathematical affine transformation, so the analysis remains understandable and reviewable to facilitate adoption by the scientific community. While currently implemented methods are simplistic and unvalidated, further development and demonstration of this framework could enable high-throughput, accurate, and accessible EDS characterization.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Affine Transformations to Correlate Experimental and Simulated EDS Spectra for Multi-Element Systems

Energy Dispersive X-ray Spectroscopy (EDS) is an essential technique for determining elemental concentrations and distributions within microstructures, critical for materials discovery, optimization, and qualification. However, most published EDS data is qualitative because current quantitative EDS analysis methods require extensive calibration and post-processing, limiting their practicality and widespread adoption. This work seeks to establish a framework for accelerated EDS characterization and spectrum analysis that can leverage ML to analyze correlations between various elemental compositions and resulting EDS spectra. The complex physics and data result in a high-dimensional problem that grows exponentially with the number of elements in the system and the complexity of the spectrum analysis. ML provides a way to compute and optimize the results of this highly dimensional problem in a flexible way to tailor it to the user’s specific needs and material system. However, the framework emphasizes transparency through a strictly mathematical affine transformation, so the analysis remains understandable and reviewable to facilitate adoption by the scientific community. While currently implemented methods are simplistic and unvalidated, further development and demonstration of this framework could enable high-throughput, accurate, and accessible EDS characterization.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Data Quality Monitoring for the Hadron Calorimeters Using Transfer Learning for Anomaly Detection

The proliferation of sensors brings an immense volume of spatio-temporal (ST) data in many domains, including monitoring, diagnostics, and prognostics applications. Data curation is a time-consuming process for a large volume of data, making it challenging and expensive to deploy data analytics platforms in new environments. Transfer learning (TL) mechanisms promise to mitigate data sparsity and model complexity by utilizing pre-trained models for a new task. Despite the triumph of TL in fields like computer vision and natural language processing, efforts on complex ST models for anomaly detection (AD) applications are limited. In this study, we present the potential of TL within the context of high-dimensional ST AD with a hybrid autoencoder architecture, incorporating convolutional, graph, and recurrent neural networks. Motivated by the need for improved model accuracy and robustness, particularly in scenarios with limited training data on systems with thousands of sensors, this research investigates the transferability of models trained on different sections of the Hadron Calorimeter of the Compact Muon Solenoid experiment at CERN. The key contributions of the study include exploring TL’s potential and limitations within the context of encoder and decoder networks, revealing insights into model initialization and training configurations that enhance performance while substantially reducing trainable parameters and mitigating data contamination effects.

47 OTHER INSTRUMENTATION

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES

Method for Pre-Conditioning a Measured Surface Height Map for Model Validation

This software allows one to up-sample or down-sample a measured surface map for model validation, not only without introducing any re-sampling errors, but also eliminating the existing measurement noise and measurement errors. Because the re-sampling of a surface map is accomplished based on the analytical expressions of Zernike-polynomials and a power spectral density model, such re-sampling does not introduce any aliasing and interpolation errors as is done by the conventional interpolation and FFT-based (fast-Fourier-transform-based) spatial-filtering method. Also, this new method automatically eliminates the measurement noise and other measurement errors such as artificial discontinuity. The developmental cycle of an optical system, such as a space telescope, includes, but is not limited to, the following two steps: (1) deriving requirements or specs on the optical quality of individual optics before they are fabricated through optical modeling and simulations, and (2) validating the optical model using the measured surface height maps after all optics are fabricated. There are a number of computational issues related to model validation, one of which is the "pre-conditioning" or pre-processing of the measured surface maps before using them in a model validation software tool. This software addresses the following issues: (1) up- or down-sampling a measured surface map to match it with the gridded data format of a model validation tool, and (2) eliminating the surface measurement noise or measurement errors such that the resulted surface height map is continuous or smoothly-varying. So far, the preferred method used for re-sampling a surface map is two-dimensional interpolation. The main problem of this method is that the same pixel can take different values when the method of interpolation is changed among the different methods such as the "nearest," "linear," "cubic," and "spline" fitting in Matlab. The conventional, FFT-based spatial filtering method used to eliminate the surface measurement noise or measurement errors can also suffer from aliasing effects. During re-sampling of a surface map, this software preserves the low spatial-frequency characteristic of a given surface map through the use of Zernike-polynomial fit coefficients, and maintains mid- and high-spatial-frequency characteristics of the given surface map by the use of a PSD model derived from the two-dimensional PSD data of the mid- and high-spatial-frequency components of the original surface map. Because this new method creates the new surface map in the desired sampling format from analytical expressions only, it does not encounter any aliasing effects and does not cause any discontinuity in the resultant surface map.

Sidick, Erkin

Mapping target signatures via partial unmixing of AVIRIS data

A complete spectral unmixing of a complicated AVIRIS scene may not always be possible or even desired. High quality data of spectrally complex areas are very high dimensional and are consequently difficult to fully unravel. Partial unmixing provides a method of solving only that fraction of the data inversion problem that directly relates to the specific goals of the investigation. Many applications of imaging spectrometry can be cast in the form of the following question: 'Are my target signatures present in the scene, and if so, how much of each target material is present in each pixel?' This is a partial unmixing problem. The number of unmixing endmembers is one greater than the number of spectrally defined target materials. The one additional endmember can be thought of as the composite of all the other scene materials, or 'everything else'. Several workers have proposed partial unmixing schemes for imaging spectrometry data, but each has significant limitations for operational application. The low probability detection methods described by Farrand and Harsanyi and the foreground-background method of Smith et al are both examples of such partial unmixing strategies. The new method presented here builds on these innovative analysis concepts, combining their different positive attributes while attempting to circumvent their limitations. This new method partially unmixes AVIRIS data, mapping apparent target abundances, in the presence of an arbitrary and unknown spectrally mixed background. It permits the target materials to be present in abundances that drive significant portions of the scene covariance. Furthermore it does not require a priori knowledge of the background material spectral signatures. The challenge is to find the proper projection of the data that hides the background variance while simultaneously maximizing the variance amongst the targets.

Boardman, Joseph W.

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry

Blueprints for Training Information Bottlenecks for Collider Analyses

Dimensionality reduction is a crucial aspect of data analysis in high energy physics, even if accompanied by information loss. Several methods, including histogram- and kernel-based analyses, are only computationally feasible for low-dimensional data. Furthermore, simulation models used in HEP can often only be validated for low-dimensional data. We provide several blueprints for using machine learning to create low-dimensional data representations (continuous event variables and discrete classification labels) for use in signal discovery and parameter estimation tasks. We also describe how to design the learned representation to facilitate a) searches with unknown model parameters and b) validation of simulation models in data control regions.

43 PARTICLE ACCELERATORS

Skin-Friction Measurements in a 3-D, Supersonic Shock-Wave/Boundary-Layer Interaction

The experimental documentation of a three-dimensional shock-wave/boundary-layer interaction in a nominal Mach 3 cylinder, aligned with the free-stream flow, and 20 deg. half-angle conical flare offset 1.27 cm from the cylinder centerline. Surface oil flow, laser light sheet illumination, and schlieren were used to document the flow topology. The data includes surface-pressure and skin-friction measurements. A laser interferometric skin friction data. Included in the skin-friction data are measurements within separated regions and three-dimensional measurements in highly-swept regions. The skin-friction data will be particularly valuable in turbulence modeling and computational fluid dynamics validation.

Wideman, J. K.

Digital Distortion Caused by Traveling- Wave-Tube Amplifiers Simulated

Future NASA missions demand increased data rates in satellite communications for near real-time transmission of large volumes of remote data. Increased data rates necessitate higher order digital modulation schemes and larger system bandwidth, which place stricter requirements on the allowable distortion caused by the high-power amplifier, or the traveling-wave-tube amplifier (TWTA). In particular, intersymbol interference caused by the TWTA becomes a major consideration for accurate data detection at the receiver. Experimentally investigating the effects of the physical TWTA on intersymbol interference would be prohibitively expensive, as it would require manufacturing numerous amplifiers in addition to acquiring the required digital hardware. Thus, an accurate computational model is essential to predict the effects of the TWTA on system-level performance when a communication system is being designed with adequate digital integrity for high data rates. A fully three-dimensional, time-dependent, TWT interaction model has been developed using the electromagnetic particle-in-cell code MAFIA (Solution of Maxwell's equations by the Finite-Integration-Algorithm). It comprehensively takes into account the effects of frequency-dependent AM (amplitude modulation)/AM and AM/PM (phase modulation) conversion, gain and phase ripple due to reflections, drive-induced oscillations, harmonic generation, intermodulation products, and backward waves. This physics-based TWT model can be used to give a direct description of the effects of the nonlinear TWT on the operational signal as a function of the physical device. Users can define arbitrary excitation functions so that higher order modulated digital signals can be used as input and that computations can directly correlate intersymbol interference with TWT parameters. Standard practice involves using communication-system-level software packages, such as SPW, to predict if adequate signal detection will be achieved. These models use a nonlinear, black-box model to represent the TWTA. The models vary in complexity, but most make several assumptions regarding the operation of the high-power amplifier. When the MAFIA TWT interaction model was used, these assumptions were found to be in significant error. In addition, digital signal performance, including intersymbol interference, was compared using direct data input into the MAFIA model and using the system-level analysis tool SPW for several higher order modulation schemes. Results show significant differences in predicted degradation between SPW and MAFIA simulations, demonstrating the significance of the TWTA approximations made in the SPW model on digital signal performance. For example, a comparison of the SPW and MAFIA output constellation diagrams for a 16-ary quadrature amplitude modulation (16-QAM) signal (data shown only for second and fourth quadrants) is shown. The upper-bound degradation was calculated from the corresponding eye diagrams. In comparison to SPW simulations, the MAFIA data resulted in a 3.6-dB larger degradation.

Kory, Carol L.

Multi-head physics-informed neural networks for learning functional priors and uncertainty quantification

In numerous applications, the integration of prior knowledge and historical information is essential, particularly for tasks requiring the solution of ordinary or partial differential equations (ODEs/PDEs) in data-sparse or noisy environments. For instance, achieving accurate solutions to time-dependent PDEs with limited initial condition measurements necessitates an effective strategy for embedding prior knowledge. Hard-parameter sharing architectures in neural networks (NNs) have demonstrated success in both traditional and scientific machine learning domains, facilitating the learning of informative representations. Here, in this study, we introduce a novel, yet efficient, method to enhance physics-informed neural networks (PINNs) by incorporating a multi-head structure that enables the learning of functional priors from both empirical data and governing physical laws. This prior information can then be used to address data sparsity and high-level noise in solving ODE/PDE problems with uncertainty quantification (UQ). The approach, termed Multi-Head PINN (MH-PINN), consists of a shared body NN and multiple head NNs, each corresponding to an individual PINN instance. Our framework for functional prior learning is carried out in two stages: (1) training the MH-PINNs to develop a shared body NN alongside multiple head NNs, and (2) employing these trained head NNs to estimate a prior distribution through a normalizing flow-based density estimator. The learned functional prior can then be applied as a regularization mechanism in deterministic contexts or as an informative prior within a Bayesian inference framework, aiding in the resolution of subsequent ODE/PDE tasks. We evaluate the efficacy of MH-PINNs across five benchmark problems, including a high-dimensional parametric PDE, all characterized by data sparsity or substantial noise levels. Our findings reveal that MH-PINNs deliver accurate solutions and robust UQ, demonstrating adaptability across a range of complex and challenging scenarios.

Bayesian inference

Polynomial chaos expansions on principal geodesic Grassmannian submanifolds for surrogate modeling and uncertainty quantification

In this work we introduce a manifold learning-based surrogate modeling framework for uncertainty quantification in high-dimensional stochastic systems. Our first goal is to perform data mining on the available simulation data to identify a set of low-dimensional (latent) descriptors that efficiently parameterize the response of the high-dimensional computational model. To this end, we employ Principal Geodesic Analysis on the Grassmann manifold of the response to identify a set of disjoint principal geodesic submanifolds, of possibly different dimension, that captures the variation in the data. Since operations on the Grassmann require the data to be concentrated, we propose an adaptive algorithm based on Riemannian K-means and the minimization of the sample Fréchet variance on the Grassmann manifold to identify “local” principal geodesic submanifolds that represent different system behavior across the parameter space. Polynomial chaos expansion is then used to construct a mapping between the random input parameters and the projection of the response on these local principal geodesic submanifolds. Here, the method is demonstrated on four test cases, a toy-example that involves points on a hypersphere, a Lotka-Volterra dynamical system, a continuous-flow stirred-tank chemical reactor system, and a two-dimensional Rayleigh-Bénard convection problem.

42 ENGINEERING