Search NASA⌕ Search

SEARCH · Search NASA

Results for “high dimensional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Immersive Visualization for Scientific Data Analysis

We will present the use of immersive visualization at the National Renewable Energy Laboratory (NREL), showcasing how immersive visualization is advancing scientific research and engineering practices and transforming our day-to-day operations. We are leveraging immersive visualization to support scientific discovery and engineering in various domains, including material design, computational fluid dynamics, immersive analytics, grid modernization, digital twins, and situated visualization. We have observed several benefits across four key areas: enhanced spatial judgments, improved understanding through interaction, increased capacity to embed high-dimensional data, and improved collaboration.

immersive analytics↗

Time projection chamber for GADGET II

The established Gaseous Detector with Germanium Tagging (GADGET) detection system is used to measure weak, low-energy 𝛽-delayed proton decays. It consists of the Gaseous Proton Detector equipped with a MICROMEGAS (MM) readout to detect protons and other charged particles calorimetrically, surrounded by the Segmented Germanium Array (SeGA) for high-resolution detection of prompt 𝛾 rays. To upgrade GADGET's Proton Detector to operate as a compact time projection chamber (TPC) for the detection, three-dimensional imaging and identification of low-energy 𝛽-delayed single- and multiparticle emissions mainly of interest to astrophysical studies. A new high granularity MM board with 1024 pads has been designed, fabricated, installed, and tested. A high-density data acquisition system based on generic electronics for TPCs (GET) has been installed and optimized to record and process the gas avalanche signals collected on the readout pads. The TPC's performance has been tested using a 220 Rn 𝛼-particle source and cosmic-ray muons. In addition, decay events in the TPC have been simulated by adapting the attpcroot data analysis framework. Furthermore, a novel application of two-dimensional convolutional neural networks for GADGET II event classification is introduced. The optimization of data throughput is also addressed. The GADGET II TPC is capable of detecting and identifying 𝛼 particles as well as measuring their track direction, range, and energy. The extracted energy resolution of the GADGET II TPC using P10 gas is about 5.4% at 6.288 MeV ( 220 Rn 𝛼 events), computed using charge integration. Based on a systematic simulation study, we estimated the detection efficiency of the GADGET II TPC for protons and 𝛼 particles, respectively. It has also been demonstrated that the GADGET II TPC is capable of tracking minimum-ionizing particles (i.e., cosmic-ray muons). From these measurements, the electron drift velocity was measured under typical operating conditions. In addition to being one of the first generation of micropattern gaseous detectors (MPGDs) to utilize a resistive anode applied to low-energy nuclear physics, the GADGET II TPC will also be the first TPC surrounded by a high-efficiency array of high-purity germanium 𝛾-ray detectors. As a result, the TPC of GADGET II has been designed, fabricated, and tested and is ready for operation at the Facility for Rare Isotope Beams for radioactive-beam-line experiments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Physics-coupled data-driven design of high-temperature alloys

We present a materials design loop, which streamlines physics-coupled machine learning (ML) surrogate models to discover new alloy chemistries with improved properties. The efficacy is demonstrated by discovering a high-temperature alumina-forming austenitic (AFA) stainless steel with enhanced creep, followed by experimental validation. The ML models have been trained using a well-curated, highly consistent experimental dataset augmented with synthetic microstructural features from a computational thermodynamic approach. We have populated a large number of hypothetical AFA alloys to explore the high-dimensional composition space and have predicted their creep properties by providing the same synthetic input features obtained from the trained ML models. Uncertainties from the ML training were taken as thresholds for truncating predicted results to identify alloys with improved or deteriorated creep. Individual elemental compositions have been determined via probability density distribution analysis from the group of alloys at the top and bottom of the predicted creep values for further virtual and experimental validations. In conclusion, we anticipate that this workflow can be applied to screen desired conditions, such as chemistry and processing parameters, in high-dimensional space through physics-guided data analytics.

Alloy design↗

Efficient lattice QCD computation of radiative-leptonic-decay form factors at multiple positive and negative photon virtualities

In previous work [D. Giusti, Methods for high-precision determinations of radiative-leptonic decay form factors using lattice QCD, Phys. Rev. D 107, 074507 (2023)], we showed that form factors for radiative leptonic decays of pseudoscalar mesons can be determined efficiently and with high precision from lattice QCD using the “three-dimensional (3D) method,” in which three-point functions are computed for all values of the current insertion time and the time integral is performed at the data-analysis stage. Here, we demonstrate another benefit of the 3D method: the form factors can be extracted for any number of nonzero photon virtualities from the same three-point functions at no extra cost. We present results for the $D_s → ℓνγ*$ vector form factor as a function of photon energy and photon virtuality, for both positive and negative virtuality, for a single ensemble with 340 MeV pion mass and 0.11 fm lattice spacing. In our analysis, we separately consider the two different time orderings and the different quark flavors in the electromagnetic current. We discuss in detail the behavior of the unwanted exponentials contributing to the three-point functions, as well as the choice of fit models and fit ranges used to remove them for various values of the virtuality. While positive photon virtuality is relevant for decays to multiple charged leptons, negative photon virtuality suppresses soft contributions and is of interest in QCD-factorization studies of the form factors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Analysis of streaked images of x-ray self-emission in laser-driven spherical implosions

Imaging of x-ray self-emission provides a powerful in situ measurement of the spatial and temporal evolution of high-energy-density plasmas. However, interpretation of these measurements requires detailed understanding of the data-generating process. This work presents a case study in the interpretation of x-ray self-emission data for the specific application of streaked one-dimensional slit imaging of spherical laser-driven implosions. A comprehensive generative model of the streaked slit-imaging diagnostic is developed including detailed treatments of the radiation transfer, photometrics, and photostatistics associated with the measurement. The model is used to generate realistic synthetic streaked images and to analyze experimental streaked images to extract important physical quantities of interest. An example analysis of streaked images from implosion experiments on the OMEGA laser is presented, where the model developed in this work is used to constrain the trajectory and peak velocity of the implosion using Bayesian inference.

Bayesian inference↗

Thermodynamics of the dipole-octupole pyrochlore magnet Ce 2 Hf 2 O 7 in applied magnetic fields

The recently discovered dipole-octupole pyrochlore magnet Ce 2 Hf 2 O 7 is a promising three-dimensional quantum spin liquid candidate which shows no signs of ordering at low temperature. Here we investigate the thermodynamic response to magnetic fields applied along the global [110] direction using specific heat measurements and fits using numerical methods, and solve the corresponding magnetic structure using neutron diffraction. Specific heat data in moderate fields are reproduced well, however, at high fields the agreement is not satisfactory. We especially observe a two-step release of entropy, a finding that demands a review of both theory and experiment. We address it within the framework of three possible scenarios, including an analysis of the crystal field Hamiltonian not restricted to the two-dimensional single-ion doublet subspace. We conclusively rule out two of these scenarios and find qualitative agreement with a simple model of field misalignment with respect to the crystalline direction. As a result, we discuss the implications of our findings for [111] applied fields and for future experiments on Ce 2 Hf 2 O 7 and its sister compounds.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Scalability Analysis of Quantum Models for Stress and Emotion Detection

Stress and emotion detection from high-dimensional physiological signals is a challenging task, particularly when aiming for accurate classification across diverse behavioral states. Quantum machine learning (QML) is promising for modeling such high-dimensional data, but scalability is limited by qubit resources and the exponential cost of classical statevector simulation. This work studies the scalability of quantum support vector machines (QSVMs) for binary stress detection and three-class emotion recognition (Negative/Neutral/Positive) under varying qubit counts and angle-encoding strategies. We also present a comparison study with one-feature-per-qubit (1:1) and two-features-per-qubit (2:1) mappings. Experiments are executed on HPC infrastructure using NVIDIA CUDA-Q to evaluate performance, variance, and class-dependent separability at higher-qubit setups. Results show that larger Hilbert spaces can improve peak accuracy but may increase instability. At the same time, dense 2:1 encoding yields more consistent stress detection performance. For emotion recognition, scaling improves discrimination for classes like Negative and Positive more than Neutral. We find that effective QML scaling is task-dependent and benefits more from encoding design than simply increasing qubit count.

Onim, Md. Saif Hassan [University of Tennessee, Kn↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Efficient mapping between void shapes and stress fields using Deep Convolutional Neural Networks with sparse data

Establishing fast and accurate structure-to-property relationships is an important component in the design and discovery of advanced materials. Physics-based simulation models like the finite element method (FEM) are often used to predict deformation, stress, and strain fields as a function of material microstructure in material and structural systems. Such models may be computationally expensive and time intensive if the underlying physics of the system is complex. This limits their application to solve inverse design problems and identify structures that maximize performance. In such scenarios, surrogate models are employed to make the forward mapping computationally efficient to evaluate. However, the high dimensionality of the input microstructure and the output field of interest often renders such surrogate models inefficient, especially when dealing with sparse data. Deep convolutional neural network (CNN) based surrogate models have shown great promise in handling such high-dimensional problems. In this paper, a single ellipsoidal void structure under a uniaxial tensile load represented by a linear elastic, high-dimensional and expensive-to-query, FEM model. We consider two deep CNN architectures, a modified convolutional autoencoder framework with a fully connected bottleneck and a UNet CNN, and compare their accuracy in predicting the von Mises stress field for any given input void shape in the FEM model. Additionally, a sensitivity analysis study is performed using the two approaches, where the variation in the prediction accuracy on unseen test data is studied through numerical experiments by varying the number of training samples from 20 to 100.

surrogate modeling; convolutional neural networks;↗

Direct mechanistic connection between acoustic signals and melt pool morphology during laser powder bed fusion

Various nondestructive diagnostic techniques have been proposed for in situ process monitoring of laser powder bed fusion (LPBF), including melt pool pyrometry, whole-layer optical imaging, acoustic emission, atomic emission spectroscopy, high speed melt pool imaging, and thermionic emission. Correlations between these in situ monitoring signals and defect formation have been demonstrated with acoustic signals having been shown to predict pore formation with especially high confidence in recent machine learning studies. Here, in this work, time-resolved acoustic data are collected in both the conduction and keyhole welding regimes of LPBF-processed Ti-6Al-4V alloy. A non-dimensionalized Strouhal number analysis, used in whistle aeroacoustics, is applied to demonstrate that the acoustic signals recorded in the keyhole regimes can be directly associated with the vapor depression morphology. This mechanistic understanding developed from whistle aeroacoustics shows that acoustic monitoring during the LPBF process can provide a direct probe into the vapor depression dynamics and defect occurrence, especially in the keyhole regimes relevant to printing and defect formation.

36 MATERIALS SCIENCE↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

An efficient hybrid downscaling framework to estimate high-resolution river hydrodynamics

Flow depth and velocity are the most important hydrodynamic variables that govern various river functions, including water resources, navigation, sediment transport, and biogeochemical cycling. Existing high-resolution flow depth simulations rely on either computationally expensive river hydrodynamic models (RHMs) or data-driven models with formidable training costs, whereas data-driven modeling of flow velocity has rarely been explored. Here, using the hybrid Low-fidelity, Spatial analysis, and Gaussian process learning (LSG) model, we developed a downscaling approach to construct high-resolution flow depth and velocity from a two-dimensional (2-D) RHM simulation at coarse resolution. The LSG models were trained and tested in an urban watershed in Houston using two different hurricane-driven flood events. The high-resolution (as fine as 30 m resolution) and low-resolution (mostly 1000 m resolution) meshes include 664 724 and 14 536 grid cells, respectively. The results showed that through downscaling, the simulation errors were reduced to less than one-fourth and one-third of the errors of the low-resolution 2-D RHM for flow depth and velocity, respectively. Our analysis further revealed that the dominant uncertainty sources of the downscaled hydrodynamics are different, with flow velocity dominated by the dimensionality reduction error, which we reduced by using a regionalized training procedure. The downscaling approach achieves an 84-fold acceleration in computational time compared to the high-resolution 2-D RHM, making high-fidelity ensemble flood modeling feasible. More importantly, the developed method provides an opportunity to couple large-scale hydrodynamical processes with local physical, chemical, and biological processes in river models.

Tan, Zeli [Pacific Northwest National Laboratory (↗

A Unified Analytical Method Greenness Score ( uAMGS ) Quantifies How Microscopic Imaging Is Greener Than Conventional Liquid Chromatography

Green chemistry is a set of principles for assessing, developing, and implementing methods that are safer, more efficient, and less detrimental to the environment. The analytical method greenness score (AMGS) is one of many metrics that attempt to evaluate traditional liquid chromatography (LC) based on the energy consumption of the instrument and the safety, health risks, and environmental impact of the solvents employed. Unfortunately, in practice, the AMGS is primarily focused on traditional separation methods in the pharmaceutical industry and is not amenable to cutting-edge separation science, including miniaturization. To broaden this scope, the unified Analytical Method Greenness Score (uAMGS) is presented here, which clarifies and expands on the underlying mathematics and incorporates both dimensional and uncertainty analysis, enabling its application to a broader range of analytical techniques. The uAMGS is used to compare the greenness of two distinct methods: single-molecule microscopy (SMM) and high-performance liquid chromatography (HPLC), which were used to collect equivalent data. uAMGS determines that SMM is significantly greener than HPLC due primarily to decreased solvent consumption. Overall, the uAMGS should allow chemists ranging from undergraduates to industrial PhDs to assess the greenness of a wide range of separations.

chemical separations↗

Facial Named Entity Recognition by Attention-Based Graph Convolutional Neural Network

In the realm of facial recognition and analysis, the ability to accurately cluster large datasets of facial images stands as a cornerstone for various applications, ranging from security surveillance to user biometric identification. This project evolves a novel approach to facial data clustering by embedding facial images into a high-dimensional vector space using an advanced embedding model trained on separate data and assumes a graph-like structure on the high-dimensional vectors. We find our method works significantly better than common shallow methods.

97 MATHEMATICS AND COMPUTING↗

DESI DR1 Ly α 1D power spectrum: Validation of estimators

The Data Release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI) is the largest sample to date for small-scale Lyα forest cosmology, accessed through its one-dimensional power spectrum (P 1D ). The Lyα forest P 1D is extracted from quasar spectra that are highly inhomogeneous (both in wavelength and between quasars) in noise properties due to intrinsic properties of the quasar, atmospheric and astrophysical contamination, and also sensitive to low-level details of the spectral extraction pipeline. We employ two estimators in DR1 analysis to measure P 1D : the optimal estimator and the fast Fourier transform (FFT) estimator. To ensure robustness of our DR1 measurements, we validate these two power spectrum and covariance matrix estimation methodologies against the challenging aspects of the data. First, using a set of 20 synthetic 1D realizations of DR1, we derive the masking bias corrections needed for the FFT estimator and the continuum fitting bias needed for both estimators. We demonstrate that both estimators, including their covariances, are unbiased with these corrections using the Kolmogorov-Smirnov test. Second, we substantially extend our previous suite of CCD image simulations to include 675,000 quasars, allowing us to accurately quantify the pipeline's performance. This set of simulations reveals biases at the highest k values, corresponding to a resolution error of a few percent. We base the resolution systematics error budget of DR1 P 1D on these values, but do not derive corrections from them since the simulation fidelity is insufficient for precise corrections.

Lyman alpha forest↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

A deep learning approach to fast analysis of collective Thomson scattering spectra

Fast analysis of collective Thomson scattering ion acoustic wave features using a deep convolutional neural network model is presented. The network was trained from spectra to predict the plasma parameters, including ion velocities, population fractions, and ion and electron temperatures. A fully kinetic particle-in-cell simulation was used to model a laboratory astrophysics experiment and simulate a diagnostic image of the ion acoustic wave feature. Network predictions were compared with Bayesian inference of the plasma model parameters for both the simulated and experimentally measured images. Both approaches were fairly accurate predicting the simulated image and the network predictions matched a good portion of the Bayesian results for the experimentally measured image. The Bayesian approach is more robust to noise and motivates future work to train deep learning models with realistic noise. The advantage of the deep learning model is making thousands of predictions in a few hundred milliseconds, compared to a few seconds to minutes per prediction for the optimization and Bayesian approaches presented here. The results demonstrate promising capabilities of deep learning models to analyze Thomson data orders of magnitude faster than conventional methods when using the neural network for standalone analysis. If more rigorous analysis is needed, neural network predictions can be used to quickly initialize other optimization methods and increase chances of success. This is especially useful when the dataset becomes very large or highly dimensional and manually refining initial conditions for the entire dataset are no longer tractable.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean↗