Search NASASearch

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine Learning Automation Pipeline

Machine Learning Automation Pipeline (MLAP) is a package to perform machine learning (ML) analysis in a step by step manner, starting with data extraction until analysis and prediction. The scripts provide the users option to chose an action such as "Extract", "Prep", and "Train" and numerous cases can be launched with just a single command. The inputs for each case are provided using a JSON file. The simulation results of several cases can be assessed using an automated process and analyzed for various metrics pertinent to ML analysis.

Jha, Pankaj

Hacking Kilometer-Scale Models: A Participative Model for Climate Information

In May 2025, nearly 700 participants from all around the world coalesced at 10 regional nodes and a few satellite nodes to take part in a global hackathon of kilometer-scale (horizontal grid spacing < 10 km) regional and global Earth system models. Exciting science is emerging from these efforts, ranging across novel model analysis, new ways of integrating with satellite data, and emulation with machine learning. New technologies were trialed that enable the community to work in new and complementary ways to democratize access to global information at a local scale from a set of the world’s highest-resolution climate models. The hackathon demonstrated how exascale data can be organized to be accessible to anyone. Fundamentally, the community could apply these techniques and technologies to move toward more participative models for coproduction and delivery of diverse sources of climate information for climate scientists and citizens alike.

Climate models

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL

Cyote-attack Chain Estimator

Attack Chain Estimator (ACE) Application Overview The Attack Chain Estimator (ACE) Application is a sophisticated tool designed for the ingestion, classification, sequencing, and enrichment of cybersecurity threat reports. This application leverages advanced machine learning models and extensive historical data to provide comprehensive insights into cyber threats, specifically targeting Industrial Control Systems (ICS). Purpose The primary functions of the ACE Application include: Ingestion of Cybersecurity Threat Reporting: Capable of ingesting text-based threat reports in markdown or text file format. Supports ingestion of structured data from other sources in STIX/JSON format. Classification of Report’s Text-Based Events: Utilizes a DeBERTa classifier, specifically trained on cybersecurity data, to map the events to MITRE ATT&CK for ICS Tactics and Techniques. Classification is performed using multiple Jupyter notebooks and machine learning workflows hosted as FastAPI microservices: regex_data deberta_base_35_train_hft_classifier_mlflow.ipynb hft_regex_classifier_mlflow.ipynb param_train_hft_classifier_mlflow.ipynb regex_tactic_tech.ipynb Ordering of Tactics, Techniques, and Observable Events: Sequences the identified tactics, techniques, and events to form a coherent attack chain. Enrichment with Historical Attack Chain Details: Enhances the attack chain with details from historical attacks using a Markov model developed from CyOTE Precursor Analysis Report data. The Markov model is available as a FastAPI endpoint for seamless integration. Enrichment with Adversary Emulation Capabilities Data: Integrates adversary emulation capabilities data using MITRE Caldera for OT adversary abilities UUIDs. Export of Output Files: Provides options to export the enriched attack chain in JSON or CSV formats. Routing of Output to Other Applications: Facilitates routing of output to various platforms and applications, including: Threat Intelligence Platforms COREII Scout for Threat Intelligence Analysis COREII Modeling and Simulation for Adversary Emulation Technical Description The ACE Application is an advanced cybersecurity tool designed to provide detailed threat analysis and sequence generation. It is built on a robust architecture that integrates natural language processing, machine learning, and historical data modeling. Key Components: Data Ingestion Module: Handles the input of threat reports and data from various formats, ensuring flexibility in data sources. Classification Engine: Employs DeBERTa-based classifiers hosted as FastAPI microservices to analyze and classify threat report events in accordance with the MITRE ATT&CK framework for ICS. Sequence Generator: Orders the classified events into a logical attack chain, providing clear insight into the sequence of tactics and techniques used in the threat. Enrichment Engine: Integrates historical data and adversary emulation capabilities to enhance the attack chain with valuable context and additional details. The historical data enrichment is powered by a Markov model, which is available as a FastAPI endpoint. Export and Routing Module: Facilitates the export of the enriched attack chain in multiple formats and routes the output to designated applications for further analysis or emulation.

Paul, Tony [Idaho National Laboratory (INL), Idaho

Machine learning for domain transfer between simulated and experimental 2D X-ray diffraction patterns using generative adversarial networks

X-ray diffraction (XRD) is a well-established technique for analyzing materials at an atomic level. Dynamic compression experiments (DCE), in which materials are subject to extreme pressures, can provide fundamental understanding to pressure-induced phase transitions and compression of the crystal lattice. The analysis of XRD patterns from highly compressed samples is non-trivial given the sparsity of data, high experimental costs, and the fact that the data is often marred with X-ray background and other artifacts. While accurate computational frameworks exist, they solve the forward problem—from structures and orientations to XRD patterns. Solving the inverse problem for 2D experimental diffraction patterns is currently a complex manual process of matching and comparing experimentally observed patterns to computationally generated ones. Machine learning is a promising tool for automating the matching process but often requires data-intensive architectures. Here, in this study, we use a CycleGAN to translate the domain of limited experimental data to a domain in which there is readily available simulated data. This domain shift allows data-intensive machine learning models that have only been trained on simulated XRD patterns to be used in the analysis of experiments.

Brozak, Samantha Jean [Sandia National Laboratorie

Towards revealing intrinsic vortex-core states in Fe-based superconductors through statistical discovery

Abstract In type-II superconductors, electronic states within magnetic vortices hold crucial information about the paring mechanism and can reveal non-trivial topology. While scanning tunneling microscopy/spectroscopy (STM/S) is a powerful tool for imaging superconducting vortices, it is challenging to isolate the intrinsic electronic properties from extrinsic effects like subsurface defects and disorders. Here we combine STM/STS with basic machine learning to develop a method for screening out the vortices pinned by embedded disorder in iron-based superconductors. Through a principal component analysis of large STS data within vortices, we find that the vortex-core states in Ba(Fe 0.96 Ni 0.04 ) 2 As 2 start to split into two categories at certain magnetic field strengths, reflecting vortices with and without pinning by subsurface defects or disorders. Our machine-learning analysis provides an unbiased approach to reveal intrinsic vortex-core states in novel superconductors and shed light on ongoing puzzles in the possible emergence of a Majorana zero mode.

Guo, Yueming

Learning continuous scattering length density profiles from neutron reflectivities using convolutional neural networks

Interpreting neutron reflectivity (NR) data using ad hoc multi-layer models and physics-based models provides information about spatially resolved neutron scattering length density (NSLD) profiles. Recent improvements in data acquisition systems have allowed acquiring thousands of NR curves in a couple of hours, which has led to a need for automated data analysis tools to interpret NR measurements in real-time. Here, we present a machine learning analysis workflow that uses a series of models, based on a convolutional neural network (CNN), to learn the relation between the NSLDs and the NRs, and subsequently produce continuous NSLD profiles directly from NRs. The usefulness of our CNN-based models is demonstrated by constructing NSLDs from NRs of several films containing homopolymer polyzwitterions and diblock copolymers mixed with different types of salts. Comparisons of the NSLDs with those constructed using ad hoc multi-layer models reveal a very good agreement, suggesting the potential of CNN-based models for real-time automated data analysis of NRs.

36 MATERIALS SCIENCE

Machine learning for single-ended event reconstruction in PROSPECT experiment

The Precision Reactor Oscillation and Spectrum Experiment, PROSPECT, was a segmented antineutrino detector that successfully operated at the High Flux Isotope Reactor in Oak Ridge, TN, during its 2018 run. Despite challenges with photomultiplier tube base failures affecting some segments, innovative machine learning approaches were employed to perform position and energy reconstruction, and particle classification. This work highlights the effectiveness of convolutional neural networks and graph convolutional networks in enhancing data analysis. By leveraging these techniques, a 3.3% increase in effective statistics was achieved compared to traditional methods, showcasing their potential to improve analysis performance. Furthermore, these machine learning methodologies offer promising applications for other segmented particle detectors, underscoring their versatility and impact.

47 OTHER INSTRUMENTATION

Improved Subseasonal Forecasting of Extreme Polar Vortices Using Machine Learning

Our research was focused on forecasting the position and shape of the winter stratospheric polar vortex at a subseasonal timescale of 15 days in advance. To achieve this, we employed both statistical and neural network machine learning techniques. The analysis was performed on 42 winter seasons of reanalysis data provided by NASA giving us a total of 6,342 days of data. The state of the polar vortex for determined by using geometric moments to calculate the centroid latitude and the aspect ratio of an ellipse fit onto the vortex. Timeseries for thirty additional precursors were calculated to help improve the predictive capabilities of the algorithm. Feature importance of these precursors was performed using random forest to measure the predictive importance and the ideal number of precursors. Then, using the precursors identified as important, various statistical methods were tested for predictive accuracy with random forest and nearest neighbor performing the best. An echo state network, a type of recurrent neural network that features sparsely connected hidden layer and a reduced number of trainable parameters that allows for rapid training and testing, was also implemented for the forecasting problem. Hyperparameter tuning was performed for each methods using a subset of the training data. The algorithms were trained and tuned on the first 41 years of data, then tested for accuracy on the final year. In general, the centroid latitude of the polar vortex proved easier to predict than the aspect ratio across all algorithms. Random forest outperformed other statistical forecasting algorithms overall but struggled to predict extreme values. Forecasting from echo state network suggested a strong predictive capability past 15 days, but further work is required to fully realize the potential of recurrent neural network approaches.

54 ENVIRONMENTAL SCIENCES

Updated ASME design correlations and qualification plan for powder bed fusion 316H stainless steel

This report provides an update on the Advanced Materials and Manufacturing Technologies (AMMT) program effort to qualify Laser-Powder Bed Fusion (L-PBF) 316H stainless steel for use with the ASME Boiler & Pressure Vessel Code Section III, Division 5 rules. The report summarizes progress in testing and characterizing L-PBF material at elevated temperatures by providing preliminary design data for L-PBF 316H and by comparing the elevated temperature performance of the L-PBF material to wrought and conventional fusion welded 316H. The report then updates the initial AMMT qualification plan for L-PBF 316H, originally developed in 2023, to update the accelerated qualification strategy adopted in that plan to account for the new high temperature test data. The report also explores a few methods for further accelerating the qualification process using machine learning techniques to supplement the more conventional, empirical analysis methods typically used by ASME to correlate and extrapolate time-dependent material test data.

36 MATERIALS SCIENCE

Review of Presentation by Carter Wolf

The presentation was informative on the different techniques used to study the electronic dynamics of materials on the femtosecond scale, such as ARPES and TR ARPES. The presentation also discussed in depth code analysis, and the procedure used to remove noise from noisy data via machine learning procedures. The presentation first focused on an overview of the ARPES technique, and how TR-ARPES differs from it. The presentation then explored computational techniques used for fitting ARPES and TR-ARPES data, using the pyARPES framework. After that, the presentation discussed machine learning approaches for denoising noisy data such as Noise2Noise and Noise2Self. The main issue of denoising existing ARPES data and the necessity of machine learning approaches due to the lack of clean ARPES data were very clearly articulated as a major part of the project. The implementation, training and refinement and optimization were clearly detailed, along with a full documentation of attempts that worked and attempts that did not, thus thoroughly and clearly showing the research process. Experimental techniques regarding ARPES were also briefly documented.

36 MATERIALS SCIENCE

Multi-system analysis of offshore geologic carbon storage: a review of open-source data science solutions

Geologic carbon storage projects are maturing worldwide and the footprint of deployment in the offshore is expanding. At present, there are ten projects in operation or that have been completed, more than 50 in construction and development, and dozens of characterization studies completed or underway. Offshore geologic carbon storage offers potential benefits over onshore geologic carbon storage. These offshore projects are generally remote in location, distant from population centers, and avoid complicated pore space rights while having abundant prospective storage potential. Some offshore fields targeted for carbon storage have comparatively fewer prior borehole penetrations except for areas that have been explored for petroleum production, minimizing potential issues such as pressure interference and infrastructure impacts. Yet offshore geologic carbon storage projects face distinctive technical and economic challenges, such as seafloor geohazards (e.g., seabed instability), expensive maritime transport, and meteorological-oceanographic conditions that can damage infrastructure and impact operations. Analytical capabilities and improved computational speeds have advanced engineering, earth and energy sciences in the wake of the arrival of modern data science over the last decade. These advancements have created an opportunity for integrated, multi-systems modeling approaches utilizing artificial intelligence and machine learning that are no longer limited by computational issues. Analytical tools developed alongside this advancement in data science can be leveraged to calibrate the potential advantages and challenges of carbon storage operations in the offshore. New methods and approaches that incorporate data science to analyze multiple aspects of engineered and natural systems can provide insights that complement the characterization and onsite engineering that traditional commercial and operational software addresses. These new methods and approaches can potentially improve the outcome of energy operations and carbon storage. Providing multi-system, science-driven data analytics enhances the knowledge base that offshore developers, operators, and regulatory bodies may draw from to improve offshore site selection and operational efficiency. Here, we provide a brief synopsis of geologic carbon storage efforts to date, an overview of the engineered and natural systems involved in offshore geologic carbon storage, and a review of publicly available, open-source, offshore and/or carbon storage related data- and science-driven tools developed by 2010 or later that are suitable for screening and assessing regions for offshore geologic carbon storage.

artificial intelligence

Bayesian inference analysis of jet quenching using inclusive jet and hadron suppression measurements

The JETSCAPE Collaboration reports a new determination of the jet transport parameter $\hat{q}$ in the quark-gluon plasma (QGP) using Bayesian inference, incorporating all available inclusive hadron and jet yield suppression data measured in heavy-ion collisions at the BNL Relativistic Heavy Ion Collider (RHIC) and the CERN Large Hadron Collider (LHC). This multi-observable analysis extends the previously published JETSCAPE Bayesian inference determination of $\hat{q}$, which was based solely on a selection of inclusive hadron suppression data. jetscape is a modular framework incorporating detailed dynamical models of QGP formation and evolution, and jet propagation and interaction in the QGP. Virtuality-dependent partonic energy loss in the QGP is modeled as a thermalized weakly coupled plasma, with parameters determined from Bayesian calibration using soft-sector observables. This Bayesian calibration of $\hat{q}$ utilizes active learning, a machine-learning approach, for efficient exploitation of computing resources. The experimental data included in this analysis span a broad range in collision energy and centrality, and in transverse momentum. In order to explore the systematic dependence of the extracted parameter posterior distributions, several different calibrations are reported, based on combined jet and hadron data; on jet or hadron data separately; and on restricted kinematic or centrality ranges of the jet and hadron data. Tension is observed in comparison of these variations, providing new insights into the physics of jet transport in the QGP and its theoretical formulation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Selection of Global Climate Model Data for Downscaling With Generative Machine Learning and Use in the Power Planning for Alignment of Climate and Energy Systems Project

The range of results from climate models and scenarios is important to the understanding of uncertainty in power planning analysis. A U.S. Department of Energy-funded analytic project called Power Planning for Alignment of Climate and Energy Systems is developing data and analytic methods to reflect the effects of climate change on key variables for power system planning, as part of the Grid Modernization Lab Consortium. This project will select and prepare global climate model results for use in power system planning models. A related report (Evaluation of Global Climate Models for Use in Energy Analysis) assesses the performance of various global climate models from the Coupled Model Intercomparison Project Phase 6 data archive for their historical skill with respect to energy system performance and for their future projections under multiple climate change scenarios. Building from that report, we describe the selection of a climate scenario (Shared Socioeconomic Pathway [SSP] 2-4.5) and five climate models: TaiESM1, EC-Earth3-CC, GFDL-CM4, EC-Earth3-Veg, and MPI-ESM1-2-HR. We describe the model selection criteria, which were based on the quality of the match between model results under historical conditions and on the representation of the range of future values for several variables. These results will be downscaled via an open-source generative machine learning method called Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts.

29 ENERGY PLANNING, POLICY, AND ECONOMY

BATMODS-lite [SWR-25-108]

Battery Analysis and Training Models for Optimization and Design Studies (BATMODS) is a Python package with an API for pre-built battery models. The original purpose of the package was to quickly generate synthetic data for machine learning models to train with. However, the models are generally useful for any battery simulations or analysis. BATMODS-lite includes the following: 1) A library and API for pre-built battery models 2) Kinetic/transport properties for common battery materials

Randall, Corey [National Laboratory of the Rockies

Revealing the evolution of order in materials microstructures using multi-modal computer vision

The development of high-performance materials for microelectronics, energy storage, and extreme environments depends on our ability to describe and direct property-defining microstructural order. Our present understanding is typically derived from laborious manual analysis of imaging and spectroscopy data, which is difficult to scale, challenging to reproduce, and lacks the ability to reveal latent associations needed for mechanistic models. Here, we demonstrate a multi-modal machine learning (ML) approach to describe order from electron microscopy analysis of the complex oxide La 1−x Sr x FeO 3 . We construct a hybrid pipeline based on fully and semi-supervised classification, allowing us to evaluate both the characteristics of each data modality and the value each modality adds to the ensemble. We observe distinct differences in the performance of uni- and multi-modal models, from which we draw general lessons in describing crystal order using computer vision.

36 MATERIALS SCIENCE

The Evaluation of Machine Learning Techniques for Isotope Identification Contextualized by Training and Testing Spectral Similarity

Precise gamma-ray spectral analysis is crucial in high-stakes applications, such as nuclear security. Research efforts toward implementing machine learning (ML) approaches for accurate analysis are limited by the resemblance of the training data to the testing scenarios. The underlying spectral shape of synthetic data may not perfectly reflect measured configurations, and measurement campaigns may be limited by resource constraints. Consequently, ML algorithms for isotope identification must maintain accurate classification performance under domain shifts between the training and testing data. To this end, four different classifiers (Ridge, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron) were trained on the same dataset and evaluated on twelve other datasets with varying standoff distances, shielding, and background configurations. A tailored statistical approach was introduced to quantify the similarity between the training and testing configurations, which was then related to the predictive performance. Wilcoxon signed-rank tests revealed that the OVR-wrapped XGB significantly outperformed the other algorithms, with confidence levels of 99.0% or above for the 133Ba, 60Co, 137Cs, and 152Eu sources. The findings from this work are significant as they outline techniques to promote the development of robust ML-based approaches for isotope identification.

domain adaptation

Nonlinearity of the post-spinel transition and its expression in slabs and plumes worldwide

Phase transitions in the mantle control its internal dynamics and structure. The post-spinel transition marks the upper–lower mantle boundary, where ringwoodite dissociates into bridgmanite plus ferropericlase, and its Clapeyron slope regulates mantle flow across it. This interaction has previously been assumed to have no lateral spatial variations, based on the assumption of a linear post-spinel boundary in pressure and temperature. Here we present laser-heated diamond anvil cell experiments with synchrotron X-ray diffraction to better constrain this boundary, especially at higher temperatures. Combining our data with results from the literature, and using a global analysis based on machine learning, we find a pronounced nonlinearity in the post-spinel boundary, with its slope ranging from –4 MPa/K at 2100 K, to –2 MPa/K at 1950 K, and to 0 MPa/K at 1600 K. Changes in temperature over time and space can therefore cause the post-spinel transition to have variable effects on mantle convection and the movement of subducting slabs and upwelling plumes.

58 GEOSCIENCES