Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

Towards a data-driven model of hadronization using normalizing flows

We introduce a model of hadronization based on invertible neural networks that faithfully reproduces a simplified version of the Lund string model for meson hadronization. Additionally, we introduce a new training method for normalizing flows, termed MAGIC, that improves the agreement between simulated and experimental distributions of high-level (macroscopic) observables by adjusting single-emission (microscopic) dynamics. Our results constitute an important step toward realizing a machine-learning based model of hadronization that utilizes experimental data during training. Finally, we demonstrate how a Bayesian extension to this normalizing-flow architecture can be used to provide analysis of statistical and modeling uncertainties on the generated observable distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Probabilistic data fusion and physics-informed machine learning: A new paradigm for modeling under uncertainty, and its application to accelerating the discovery of new materials

In this report we summarize the work conducted by PI Perdikaris and his group under this Early Career project DE–SC0019116 during the period of 09/01/2018 – 08/31/2023. The central aim of the work was to introduce a new paradigm for scientific data analysis that can seamlessly synthesize rigorous mathematical modeling with data of variable fidelity (e.g., measurements at multiple scales/resolutions or predictions of variable fidelity models) and multiple modalities (e.g., images, time–series, or scattered measurements). The setting we are interested in involves complex systems that are partially observed and whose dynamical behavior could be hard to model or totally unknown. The inherent uncertainty associated with this setting necessitates a departure from the classical deterministic realm of modeling and scientific computation, and, consequently, our main building blocks can no longer be crisp deterministic numbers and governing laws, but instead we must operate with probabilistic models.

97 MATHEMATICS AND COMPUTING↗

Exploring Saccharomycotina Yeast Ecology Through an Ecological Ontology Framework

Yeasts in the subphylum Saccharomycotina are found across the globe in disparate ecosystems. A major aim of yeast research is to understand the diversity and evolution of ecological traits, such as carbon metabolic breadth, insect association, and cactophily. This includes studying aspects of ecological traits like genetic architecture or association with other phenotypic traits. Genomic resources in the Saccharomycotina have grown rapidly. Ecological data, however, are still limited for many species, especially those only known from species descriptions where usually only a limited number of strains are studied. Moreover, ecological information is recorded in natural language format limiting high throughput computational analysis. To address these limitations, we developed an ontological framework for the analysis of yeast ecology. A total of 1,088 yeast strains were added to the Ontology of Yeast Environments (OYE) and analyzed in a machine-learning framework to connect genotype to ecology. This framework is flexible and can be extended to additional isolates, species, or environmental sequencing data. Widespread adoption of OYE would greatly aid the study of macroecology in the Saccharomycotina subphylum.

59 BASIC BIOLOGICAL SCIENCES↗

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L↗

Evolution and Degradation Patterns of Electrochemical Cells Based on the Analysis of Interfacial Phenomena at Li Metal Anode/Electrolyte Interfaces

In this work, we report the results of a theoretical–computational analysis of the solid electrolyte interphase (SEI) growth and degradation dynamics occurring in lithium metal batteries during cycling. We use ab initio-kinetic Monte Carlo simulations to generate a synthetic data set, which is analyzed by machine learning methods. We aim to determine: (i) how modifications in interfacial interaction energies between solid electrolyte interphase (SEI) blocks and between Li ions and SEI facets impact the Coulombic efficiency (CE) of the battery and (ii) what factors, including reactions, microscopic transport, and other interfacial events, may lead to cell performance “failure” during prolonged charge and discharge cycles, signaled as a sharp decay in the CE over cycling. The demonstration of our approach is done on a cell including a Li metal surface interfacing with a previously introduced state-of-the-art electrolyte, and the idea can be applied to any electrochemical system. Outcomes include the identification of the leading chemical, physical, and structural variables causing cell failure and relating them to the electrolyte formulation, thus paving the way to future more refined analysis and electrolyte design.

batteries↗

DONUT: physics-aware machine learning for real-time X-ray nanodiffraction analysis

Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Materials science↗

Machine-learning-informed scattering correlation analysis of sheared colloids

We have carried out theoretical analysis, Monte Carlo simulations and machine-learning analysis to quantify microscopic rearrangements of dilute dispersions of spherical colloidal particles from coherent scattering intensity. Both monodisperse and polydisperse dispersions of colloids were created and underwent a rearrangement consisting of an affine simple shear and non-affine rearrangement using the Monte Carlo method. We calculated the coherent scattering intensity of the dispersions and the correlation function of intensity before and after the rearrangement and generated a large data set of angular correlation functions for varying system parameters, including number density, polydispersity, shear strain and non-affine rearrangement. Singular value decomposition of the data set shows the feasibility of machine-learning inversion from the correlation function for the polydispersity, shear strain and non-affine rearrangement using only three parameters. A Gaussian process regressor is then trained on the data set and can retrieve the affine shear strain, non-affine rearrangement and polydispersity with relative errors of 3%, 1% and 6%, respectively. Altogether, our model provides a framework for quantitative studies of both steady and non-steady microscopic dynamics of colloidal dispersions using coherent scattering methods.

Gaussian process regression↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Machine Learning-Based Energy Estimation for Sterile Neutrino Searches in the NOνA Experiment

This dissertation presents a search for sterile neutrinos using Monte Carlo datasets and experimental data from the NOνA experiment. This work introduces novel machine learning techniques for energy reconstruction in neutral current events. A deep learning-based energy estimator was developed and integrated into the analysis framework. Using this new energy estimator, the sensitivity to sterile neutrino-induced oscillations is evaluated and presented. Furthermore, the dissertation explores potential methods for further improvement of the analysis.

43 PARTICLE ACCELERATORS↗

Predicting High‐Resolution Spatial and Spectral Features in Mass Spectrometry Imaging with Machine Learning and Multimodal Data Fusion

Recent advancements in molecular Mass Spectrometry Imaging have sparked interest in integrating high spatial resolution methods with molecular mass-spectrometry-based chemical imaging. Fusion-based algorithms have proven effective in generating high spatial-resolution molecular mass spectra. However, a significant challenge stems from the differing physical mechanisms underlying image generation and data upsampling techniques, potentially leading to discrepancies in integrated information channels. Integrating physical constraints into data processing workflows is essential to tackle this issue. In this study, we propose an innovative approach that merges data from Fourier transform ion cyclotron resonance (FTICR), time-of-flight matrix-assisted laser desorption/ionization, and time-of-flight secondary ion mass spectrometry imaging techniques. By leveraging FT-ICR's unparalleled spectral resolution and ToF-SIMS's exceptional spatial resolution, we achieve submicron spatial resolution, enabling the observation of intact molecular species with remarkable spectral precision. Canonical correlation analysis is employed to incorporate physical constraints. Through sophisticated image processing and machine learning techniques, the results of this fusion hold significant promise for advancing our comprehension of complex systems and unveiling concealed molecular intricacies.

canonical correlation analysis↗

DONUT: Physics-aware Machine Learning for Real-time X-ray Nanodiffraction Analysis

SF-25-088 Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Zhou, Tao [Argonne National Laboratory (ANL), Argo↗

Advances in geophysical forensic event monitoring

Forensic analysis of man-made, non-nuclear events (such as industrial accidents, explosion experiments and mine collapses) has become more frequent and detailed owing to advancements in geophysical monitoring. Here, in this Technical Review, we demonstrate how geophysical forensic monitoring using seismic, infrasound and hydroacoustic recordings provides insights on events in the solid earth, atmosphere and underwater. Advanced techniques, including machine-learning-based models, have been developed to detect, identify and investigate these events, providing information on location, subevents, sources and explosive yield. The increase in data availability, application of advanced methods and computation and the growth of multitechnology approaches have increased the accuracy of forensic event analysis and enabled more realistic characterization of uncertainties. For example, the 2020 Beirut explosion in Lebanon demonstrated that various seismic, acoustic and other methods could be used to estimate explosive yield (and yield uncertainties) of about 1 ktonne, providing confidence in the application of these methods to smaller events where data are available. However, forensic investigations remain largely limited to known events with identified sources. Increased access to data, sophisticated analysis methods and high-resolution earth models will improve forensic event analysis further, enabling civil and scientific applications, such as localization in the search for the lost ARA San Juan submarine.

geophysics↗

AutoEMX v1.

The invention consists in the full automation of compositional analysis of inorganic powder samples by scanning electron microscopy (SEM) with energy-dispersive X-ray spectroscopy (EDS). The measurements and analysis are controlled via python-based software, Auto-SEMEDS. Auto-SEMEDS fully automates the SEM-EDS measurements, and analyses the collected data via the use of machine-learning (ML) algorithms, which have never been used before for such scope. Auto-SEMEDS enables the identification in fully-automated fashion of the individual material phases present in a powder sample. Similar technologies, such as commercial SEM-EDS software, can automatically classify particles based on their composition, but they have significant limitations. These solutions typically provide inaccurate composition measurements and struggle to identify single phases in lab samples, where phases are often closely intermixed. In contrast, Auto-SEMEDS achieves unprecedented accuracy in composition measurements of powder samples, and furthermore leverages machine learning algorithms to effectively discern intermixed phases. Notably, while previous studies have demonstrated accurate measurements on individual particles, Auto-SEMEDS stands out by successfully analyzing mixture of different phases, a capability that has not been reported in the literature until now.

Giunto, Andrea [Lawrence Berkeley National Laborat↗

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

Spectral kernel machines with electrically tunable photodetectors

Spectral machine vision collects spectral and spatial information as three-dimensional hypercubes and digitally processes them, which causes a data bottleneck, limiting power efficiency, frame rate, and spectral-spatial resolution. This work introduces spectral kernel machines (SKMs) to overcome these bottlenecks. SKM directly compresses spectral analysis through the output photocurrent and learns from example objects to identify and classify new samples in a "sniff-and-seek" mode. We experimentally demonstrated SKMs with electrically tunable bipolar black phosphorus-molybdenum disulfide (bP-MoS2) photodiodes in the near- and mid-infrared band and silicon photoconductors in the visible band, performing versatile intelligent tasks from chemometrics to semiconductor metrology. This architecture consumed substantially less power and was more than an order of magnitude faster than existing solutions for hyperspectral image analysis, defining an intelligent imaging and sensing paradigm with intriguing possibilities.

Zhang, Dehui↗

Livewire: Automatic Annotations

Diogenes processes datasets to provide data quality metrics for the Livewire platform and creates standardized data dictionaries from data annotations. Diogenes needs data annotations that clearly outline thenformat and organization of the data. It also relies on the type, class, and unit of each data piece for comprehensive analysis, which it cannot determine independently. The Annotation Tool significantly reduces the time needed to create annotations for Diogenes by generating data annotations with the correct formatting and content. It also employs machine learning and hard-coded models to automatically annotate data class, quality type, and data units.

33 - ADVANCED PROPULSION SYSTEMS↗