Search NASA⌕ Search

SEARCH · Search NASA

Results for “High dimensional data,”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Data-driven global ocean modeling for seasonal to decadal prediction

Accurate modeling of ocean dynamics is crucial for enhancing our understanding of complex ocean circulation processes, predicting climate variability, and tackling challenges posed by climate change. Although great efforts have been made to improve traditional numerical models, predicting global ocean variability over multiyear scales remains challenging. Here, we propose ORCA-DL (Oceanic Reliable foreCAst via Deep Learning), a data-driven three-dimensional ocean model for seasonal to decadal prediction of global ocean dynamics. ORCA-DL accurately simulates the three-dimensional structure of global ocean dynamics with high physical consistency and outperforms state-of-the-art numerical models in capturing extreme events, including El Niño–Southern Oscillation and upper ocean heat waves. Moreover, ORCA-DL stably emulates ocean dynamics at decadal timescales, demonstrating its potential even for skillful decadal predictions and climate projections. Our results demonstrate the high potential of data-driven models for providing efficient and accurate global ocean modeling and prediction.

Science & Technology - Other Topics↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

Machine Learning a Simple Interpretable Short-Range Potential for Silica

A wide array of models, spanning from computationally expensive ab initio methods to a spectrum of force-field approaches, have been developed and employed to probe silica polymorphs and understand growth processes and atomic-level dynamical transitions in silica. However, the quest for a model capable of making accurate predictions with high computational efficiency for various silica polymorphs is still ongoing. Recent developments in short-range machine-learned models, such as GAP and NNPScan, have shown promise in providing reasonable descriptions of silica, but their computational cost remains high compared to force fields such as BKS which are based on simple interpretable functional forms. Here, in this study, we build on the recent success of our reinforcement learning (RL) workflow to derive a new set of optimal parameters for a promising short-range BKS-based model proposed by Soules. We use RL to navigate the eight-dimensional parameter space of the Soules potential using an experimental training data set that includes both local and global structural features from approximately 21 experimentally realized silica polymorphs, including high density phases and porous zeolites. We compare the performance of our machine-learned ML-Soules model with other high quality models including our recent machine-learned parametrization of BKS (ML-BKS), a machine-learned potential (GAP), as well as predictions of ab initio calculations with the highly fidelity SCAN functional. The ML-Soules accurately captures the relative energetic ordering of various polymorphs as well as their structural features at a significantly reduced computational expense. The ML-Soules model also reasonably captures the structure, density, and elastic constants of quartz, as well as metastable silica polymorphs. We further discuss the limitations of the Soules functional form and propose potential enhancements, including the incorporation of additional three-body terms and/or the utilization of different short-ranged functional forms to achieve greater accuracy for both global and local features in the modeling of silica while retaining low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fractal Scaling of Explosively Driven Product Gases

ABSTRACT Characterization of the interface between explosive product gases and ambient air in an explosion is a complicated task due to the turbulent mixing and inherently three‐dimensional expansion of the interface. This study aims to quantify the evolution of the interface as a temporally varying Hausdorff dimension. Two test series were conducted with Composition C‐4 charges with masses of 105 and 880 g. Imaging data were collected from the time of detonation until shock wave detachment using ultra‐high‐speed cameras. Gas cloud profiles were extracted using automated image processing algorithms, and the Hausdorff dimension of these two‐dimensional slices of the gas cloud was then estimated using boxcounting algorithms. When scaled with standard gas dynamic nondimensional scalings, the Hausdorff dimension of all explosive events appears to collapse towards a single curve. The fireball was initially nonfractal and began to develop fractal properties as the shock wave separated from the detonation products. Artificial perturbation of the charged surface had no detectable impact on the evolution of the Hausdorff dimension in the early development of the fireball outside of error, despite visible phenomenological differences in the early development of mixing on the fireball surface.

42 ENGINEERING↗

Unorthodox parallelization for Bayesian quantum state estimation

Quantum state tomography (QST) allows for the reconstruction of quantum states through measurements and some inference technique under the assumption of repeated state preparations. Bayesian inference provides a promising platform to achieve both efficient QST and accurate uncertainty quantification, yet is generally plagued by the computational limitations associated with long Markov chains. In this work, we present a novel Bayesian QST approach that leverages modern distributed parallel computer architectures to efficiently sample a D-dimensional Hilbert space. Using a parallelized preconditioned Crank–Nicholson Metropolis–Hastings algorithm, we demonstrate our approach on simulated data and experimental results from IBM Quantum systems up to four qubits, showing significant speedups through parallelization. Although highly unorthodox in pooling independent Markov chains, our method proves remarkably practical, with validation ex post facto via diagnostics like the intrachain autocorrelation time. We conclude by discussing scalability to higher-dimensional systems, offering a path toward efficient and accurate Bayesian characterization of large quantum systems.

Bayesian inference↗

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Biopolymer-Templated Titania Film Formation for Nanostructured Coatings Revealed by Machine Learning-Supported Time-Resolved Analysis

This study presents a machine learning approach to derive the film formation of biopolymer-templated titania nanostructures during spray deposition, in combination with in situ grazing-incidence small-angle X-ray scattering (GISAXS). A neural network trained on synthetic GISAXS data directly predicts domain-size distributions from experimental two-dimensional scattering patterns, capturing the full kinetics of nanostructure evolution with high temporal resolution. The predictions reveal hierarchical size distributions and periodic growth features, consistent with layer-by-layer spray deposition and validated by complementary scanning electron microscopy (SEM) imaging. Quantitative comparison with conventional parametric GISAXS fits shows good qualitative agreement, with systematic differences explained by domain-shape assumptions and resolved by applying a geometric scaling factor. Simulated SEM-like surfaces derived from neural network outputs reproduce the porous, foam-like nanoscale morphology observed experimentally, reinforcing the method’s credibility. This integrated approach enables real-time, nondestructive, statistically averaged monitoring of bulk nanostructure development in functional coatings, offering a scalable methodology to accelerate the characterization and process control of sustainably manufactured nanostructured titania films for energy-related applications such as photocatalysis and photovoltaics.

Heger, JulianEliah↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Time projection chamber for GADGET II

The established Gaseous Detector with Germanium Tagging (GADGET) detection system is used to measure weak, low-energy 𝛽-delayed proton decays. It consists of the Gaseous Proton Detector equipped with a MICROMEGAS (MM) readout to detect protons and other charged particles calorimetrically, surrounded by the Segmented Germanium Array (SeGA) for high-resolution detection of prompt 𝛾 rays. To upgrade GADGET's Proton Detector to operate as a compact time projection chamber (TPC) for the detection, three-dimensional imaging and identification of low-energy 𝛽-delayed single- and multiparticle emissions mainly of interest to astrophysical studies. A new high granularity MM board with 1024 pads has been designed, fabricated, installed, and tested. A high-density data acquisition system based on generic electronics for TPCs (GET) has been installed and optimized to record and process the gas avalanche signals collected on the readout pads. The TPC's performance has been tested using a 220 Rn 𝛼-particle source and cosmic-ray muons. In addition, decay events in the TPC have been simulated by adapting the attpcroot data analysis framework. Furthermore, a novel application of two-dimensional convolutional neural networks for GADGET II event classification is introduced. The optimization of data throughput is also addressed. The GADGET II TPC is capable of detecting and identifying 𝛼 particles as well as measuring their track direction, range, and energy. The extracted energy resolution of the GADGET II TPC using P10 gas is about 5.4% at 6.288 MeV ( 220 Rn 𝛼 events), computed using charge integration. Based on a systematic simulation study, we estimated the detection efficiency of the GADGET II TPC for protons and 𝛼 particles, respectively. It has also been demonstrated that the GADGET II TPC is capable of tracking minimum-ionizing particles (i.e., cosmic-ray muons). From these measurements, the electron drift velocity was measured under typical operating conditions. In addition to being one of the first generation of micropattern gaseous detectors (MPGDs) to utilize a resistive anode applied to low-energy nuclear physics, the GADGET II TPC will also be the first TPC surrounded by a high-efficiency array of high-purity germanium 𝛾-ray detectors. As a result, the TPC of GADGET II has been designed, fabricated, and tested and is ready for operation at the Facility for Rare Isotope Beams for radioactive-beam-line experiments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

Designing Remote Monitoring for Smart Manufacturing Facilities: Hazard Identification and Classification

This study investigates the process of hazard identification in complex manufacturing environments during the design phase, emphasizing the significance of the design process in developing designs that effectively mitigate hazards in contexts with numerous variables, such as a variety of machines, sensors, actuators, and agents. Through a mixed-methods approach, the objective of this work is to understand how the evolution of design outcomes across various stages might influence a designer’s ability to recognize both standard and novel hazards. To achieve this understanding, an experimental design task was conducted with six designers from a national lab specializing in manufacturing technologies. This approach combined qualitative and quantitative data analysis from a one-hour virtual session with participants. Findings suggest that the complexity of identifying hazards in a high-dimensional design space is challenging within a limited time frame and that the identification of hazards is significantly influenced by the stage of the design task and the initial design decisions, indicating the need for extended time and strategic initial planning in the design process to enhance hazard identification.

Ballestas, Caseysimone↗

Understanding the interplay between pilot fuel mixing and auto-ignition chemistry in hydrogen-enriched environment

The diesel-piloted dual-fuel compression ignition combustion strategy is well-suited to accelerate the decarbonization of transportation by adopting hydrogen as a renewable energy carrier into the existing internal combustion engine with minimal engine modifications. Despite the simplicity of engine modification, many questions remain unanswered regarding the optimal pilot injection strategy for reliable ignition with minimum pilot fuel consumption. The present study uses a single-cylinder heavy-duty optical engine to explore the phenomenology and underlying mechanisms governing the pilot fuel ignition and the subsequent combustion of a premixed hydrogen-air charge. The engine is operated in a dual-fuel mode with hydrogen premixed into the engine intake charge with a direct pilot injection of n-heptane as a diesel pilot fuel surrogate. Optical diagnostics used to visualize in-cylinder combustion phenomena include high-speed IR imaging of the pilot fuel spray evolution as well as high-speed HCHO* and OH* chemiluminescence as indicators of low-temperature and high-temperature heat release, respectively. Three pilot injection strategies are compared to explore the effects of pilot fuel mass, injection pressure, and injection duration on the probability and repeatability of successful ignition. The thermodynamic and imaging data analysis supported by zero-dimensional chemical kinetics simulations revealed a complex interplay between the physical and chemical processes governing the pilot fuel ignition process in a hydrogen containing charge. Hydrogen strongly inhibits the ignition of pilot fuel mixtures and therefore requires longer injection duration to create zones with sufficiently high pilot fuel concentration for successful ignition. Results show that ignition typically tends to rely on stochastic pockets with high pilot fuel concentration, which results in poor repeatability of combustion and frequent misfiring. In conclusion, this work has improved the understanding on how the unique chemical properties of hydrogen pose a challenge for maximization of hydrogen’s energy share in hydrogen dual-fuel engines and highlights a potential mitigation pathway.

33 ADVANCED PROPULSION SYSTEMS↗

Ab Initio-Based Bond Order Potential for Arsenene Polymorphs Developed via Hierarchical Reinforcement Learning

Arsenene, a less-explored two-dimensional material, holds the potential for applications in wearable electronics, memory devices, and quantum systems. This study introduces a bond-order potential model with Tersoff formalism, the ML-Tersoff, which leverages multireward hierarchical reinforcement learning (RL), trained on an ab initio data set. This data set covers a spectrum of properties for arsenene polymorphs, enhancing our understanding of its mechanical and thermal behaviors without the complexities of traditional models requiring multiple parameter sets. Our RL strategy utilizes decision trees coupled with a hierarchical reward strategy to accelerate convergence in high-dimensional continuous search spaces. Unlike the Stillinger-Weber approach, which demands separate formalisms for buckled and puckered forms, the ML-Tersoff model concurrently captures multiple properties of the two polymorphs by effectively representing the local environment, thereby avoiding the need for different atomic types. Here, we apply the ML model to understand the mechanical and thermal properties of the arsenene polymorphs and nanostructures. We observe an inverse relationship between the critical strain and temperature in arsenene. Thermal conductivity calculations in nanosheets show good agreement with ab initio data, reflecting a decrease in thermal conductivity attributable to increased anharmonic effects at higher temperatures. We also apply the model to predict the thermal behavior of arsenene nanotubes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An In Situ , Automated High-Explosives Aging Method Utilizing Two-Dimensional Gas Chromatography–Mass Spectrometry

Understanding chemical changes that occur in high explosives as they age is of great importance to the safe employment and storage of these compounds. Traditional methods of aging high explosives even under accelerated aging conditions are time intensive with durations on the order of months to years. The nature of traditional aging analyses reduces each sample to a snapshot data point often separated widely in time, requiring many assumptions as to how the degradation products develop. Further complicating matters, several analytical techniques are typically employed for each sample analysis in order to ascertain an entire picture of the decomposition pathways. To address these shortcomings with existing methods, a new method of accelerated aging of high explosives utilizing comprehensive two-dimensional gas chromatography coupled to high-resolution mass spectrometry (GC × GC-HRMS) was developed using 2,4,6,8,10,12-hexanitro-2,4,6,8,10,12-hexaazaisowurtzitane (CL-20) as a model compound for method development. This in situ automated method reduces the time scale of aging to a matter of hours using the inlet of the GC × GC as the aging vessel. GC × GC in combination with HRMS allowed for the collection of both evolved gases and other decomposition products produced during the entire aging process in real time with HRMS providing far greater certainty in identification of explosives aging products. Additionally, this method allowed for a higher throughput of samples with greatly simplified sample preparation. Chemometric analysis of the GC × GC-HRMS data set via the alteration analysis (ALA) enabled discovery of statistically significant chemical changes providing insight into the variation of decomposition pathways with varying aging temperatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Unsupervised discovery of extreme weather events using universal representations of emergent organization

Spontaneous self-organization is ubiquitous in systems far from thermodynamic equilibrium. While organized structures that emerge dominate transport properties, universal representations that identify and describe these key objects remain elusive. Here, we introduce a theoretically grounded framework for describing emergent organization that, via data-driven algorithms, is constructive in practice. Its building blocks are spacetime lightcones that embody how information propagates across a system through local interactions. We show that predictive equivalence classes of lightcones—local causal states—capture organized behaviors in complex spatiotemporal systems. Employing an unsupervised physics-informed machine learning algorithm and a high-performance computing implementation, we demonstrate automatically discovering organized structures in two real-world domain science problems. We show that local causal states identify vortices and track their power-law decay behavior in two-dimensional fluid turbulence. We then show how to detect and track familiar extreme weather events—hurricanes and atmospheric rivers—and discover other novel structures associated with precipitation extremes in high-resolution climate data at the grid-cell level.

Rupe, Adam [Pacific Northwest National Laboratory ↗

Scaling kinetic Monte-Carlo simulations of grain growth with combined convolutional and graph neural networks

Graph neural networks (GNN) have emerged as a promising machine learning method for microstructure simulations such as grain growth. However, accurate modeling of realistic grain boundary networks requires large simulation cells, which GNN has difficulty scaling up to. To alleviate the computational costs and memory footprint of GNN, we suggest a hybrid architecture combining a convolutional neural network (CNN) based bijective autoencoder to compress the spatial dimensions, and a GNN that evolves the microstructure in the latent space of reduced spatial sizes. Our results demonstrate that the new design significantly reduces computational costs with using fewer message passing layer (from 12 down to 3) compared with GNN alone. The reduction in computational cost becomes more pronounced as the spatial size increases, indicating strong computational scalability. For the largest mesh evaluated (160 3 ), our method reduces memory usage and runtime in inference by 117× and 115×, respectively, compared with GNN-only baseline. More importantly, it shows higher accuracy and stronger spatiotemporal capability than the GNN-only baseline, especially in long-term testing. Such combination of scalability and accuracy is essential for simulating realistic material microstructures over extended time scales. The improvements can be attributed to the bijective autoencoder’s ability to compress information losslessly from spatial domain into a high dimensional feature space, thereby producing more expressive latent features for the GNN to learn from, while also contributing its own spatiotemporal modeling capability. Training data are generated from stochastic grain growth simulations, providing realistic variability for learning robust microstructure evolution. Comprehensive system validation confirms that the model is accurate, robust, and scalable.

36 MATERIALS SCIENCE↗

Light in the dark forest. Part I. An efficient optimal estimator for 3D Lyman-alpha forest power spectrum

The highly anisotropic nature of the Lyman-alpha (Lyα) forest data introduces a complex survey window function that complicates the measurement of the three-dimensional power spectrum ( P 3D ). In this paper, we present the first fully optimal estimator for P 3D , which exactly deconvolves the survey window function and marginalizes contaminated modes that distort the power spectrum. Our approach adapts optimal estimator techniques developed for the 2D cosmic microwave background data to the 3D case. To achieve computational feasibility, we employ the conjugate gradient method and implement the P 3 M formalism to handle large-scale and small-scale operations separately and efficiently. We validate our estimator using Monte Carlo mocks and Gaussian simulations, demonstrating its accuracy and computational efficiency. We confirm that mode marginalization eliminates distortions arising from quasar continuum errors and delivers robust power spectrum estimation, though it also inflates errors at large scales. This first implementation works in the flat-sky case; we discuss the remaining steps needed to generalize it to the curved-sky case. This formalism offers a foundation for the Lyα forest P 3D measurements and a new path toward cosmological constraints from the Lyα forest data.

Lyman alpha forest↗