Search NASA⌕ Search

SEARCH · Search NASA

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

The ab initio non-crystalline structure database: empowering machine learning to decode diffusivity

Non-crystalline materials exhibit unique properties that make them suitable for various applications in science and technology, ranging from optical and electronic devices and solid-state batteries to protective coatings. However, data-driven exploration and design of non-crystalline materials is hampered by the absence of a comprehensive database covering a broad chemical space. In this work, we present the largest computed non-crystalline structure database to date, generated from systematic and accurate ab initio molecular dynamics (AIMD) calculations. We also show how the database can be used in simple machine-learning models to connect properties to composition and structure, here specifically targeting ionic conductivity. These models predict the Li-ion diffusivity with speed and accuracy, offering a cost-effective alternative to expensive density functional theory (DFT) calculations. Furthermore, the process of computational quenching non-crystalline structures provides a unique sampling of out-of-equilibrium structures, energies, and force landscape, and we anticipate that the corresponding trajectories will inform future work in universal machine learning potentials, impacting design beyond that of non-crystalline materials. In addition, combining diffusion trajectories from our dataset with models that predict liquidus viscosity and melting temperature could be utilized to develop models for predicting glass-forming ability.

36 MATERIALS SCIENCE↗

CoRE MOF DB: A curated experimental metal-organic framework database with machine-learned properties for integrated material-process screening

Here, we present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.

CoRE MOF database↗

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

A traffic accident dataset for Chattanooga, Tennessee

This publication presents an annotated accident dataset which fuses traffic data from radar detection sensors, weather condition data, and light condition data with traffic accident data (as illustrated in Fig. 1) in a format that is easy to process using machine learning tools, databases, or data workflows. The purpose of this data is to analyze, predict, and detect traffic patterns when accidents occur. Each file contains a timeseries of traffic speeds, flows, and occupancies at the sensor nearest to the accident, as well as 5 neighboring sensors upstream and downstream. It also contains information about the accident type, date, and time. In addition to the accident data, we provide baseline data for typical traffic patterns during a given time of day. Overall, the dataset contains 6 months of annotated traffic data from November 2020 to April 2021. During this timeframe, and 361 accidents occurred in the monitored area around Chattanooga, Tennessee. This dataset served as the basis for a study on topology-aware automated accident detection for a companion publication [1].

97 MATHEMATICS AND COMPUTING↗

Machine learning-enabled discovery of ionic liquid–solvent electrolytes exhibiting high ionic conductivity

Ionic liquids (ILs), which are a class of materials with versatile nature and growing popularity, are facing impediments toward widespread usage as electrolytes due to various factors such as low ionic conductivity, high viscosity, high market price etc. One of the ways these limitations can be addressed is by mixing ILs with a molecular solvent. In a combinatorial sense, there exists an immense number of specific IL–solvent combinations. An exhaustive experimental or even simulation-based investigation of the chemical space spanned by such combinations can be extremely time-consuming, expensive, and nearly impossible. An alternative approach is to employ machine learning-based models developed from available databases. Although there exists prior literature that integrates machine learning to investigate mixtures of specific solvents with ILs, these models lack generalization necessitating development of a large number of ML models to handle various solvents. To remedy this shortcoming, as a part of designing green electrolytes with high ionic conductivity that can have potential applications in next-generation batteries and solar cells, this work aims to develop a unified machine learning model to predict ionic conductivity of any IL–solvent mixture system. In this regard, three models, namely, Random Forest, extreme gradient boosting (XGBoost), and artificial neural network (ANN) were formulated using the NIST ILThermo database. The dataset contained 549 unique ionic liquids from 16 cation families and 81 unique solvents, representing a total of 23 712 datapoints. SHAPLEY additive explanation (SHAP) method was used to assess the impact of various features on model prediction and their significance was compared with literature to gain physical insight about the model behavior. Finally, using the developed models, approximately 2.5 million IL–solvent mixtures at five different compositions were screened at room temperature. The high-throughput screening yielded nearly 19 000 IL–solvent mixtures for which ionic conductivity was found to exceed the ionic conductivity of conventional Li-ion battery electrolyte.

25 ENERGY STORAGE↗

Design, Control and Application of Next Generation Qubits

Design, Control and Application of Next Generation Qubits Arun Bansil, Northeastern University (Principal Investigator) Claudio Chamon, Boston University (Co-Investigator) Adrian Feiguin, Northeastern University (Co-Investigator) Liang Fu, MIT (Co-Investigator) Eduardo Mucciolo, Univ. of Central Florida (Co-Investigator) Qimin Yan, Temple University (Co-Investigator) The quest for developing technologies for manipulating and storing information quantum mechanically is currently led by approaches that include Josephson-junctions, ion-traps, and qubits generated by defect spins in solids. Topological qubits, however, are inherently more robust to decoherence by environmental effects, and should be able to sprint ahead once practical barriers have been overcome. At the present stage of the development of the field, it is important to explore a variety of architectures and materials beyond the conventional paradigms in order to seed breakthroughs toward building a scalable quantum computer. Our comprehensive theoretical research program involved four interconnected thrusts as follows. • A materials discovery effort in two-dimensional compounds in search of materials to support Majorana zero modes and defect structures suitable as qubits. • Exploration of architectures for topological quantum computation by investigating both superconducting Majorana qubits, and robust platforms for braiding with new “meta-materials” built of arrays of Majorana qubits. • Investigation of properties of hybrid metal-organic qubits based on transition-metal centers in graphene, and molecular crystals of polyaromatic complexes with embedded transition-metal atoms. • Development of tensor-network and semiclassical approaches to study decoherence in the presence of random and dispersive spin baths, and NV centers in diamond. The full spectrum of theoretical and numerical approaches was used to address the goals of this project including first-principles, density-matrix-renormalization group, tensor networks, and data-driven high-throughput approaches using materials database and machine-learning.

36 MATERIALS SCIENCE↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bia. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Ref. [2]. The prerequisite for applying machine learning techniques is casting the metadata into a format that can be parsed by the algorithm. This step might seem trivial but requires to find a unique language where metadata that carry the same physics meaning across several experiments must have the same identifier. One example is, for instance, the neutron detector. As seen in Figure 1, the machine learning code identified the use of 6 Li detectors as being related to bias in some datasets of the AIACHNE 252 Cf PFNS experimental database. In fact, here are several experiments that used neutron detectors containing 6Li in the database, for instance for the example below. EXFOR format has a unique keywords describing detectors such as “SCIN” or “GLASD”. One may think that these keywords are already sufficient descriptors for ML to uniquely find an issue. However, “SCIN” (used for [3, 4]) and “GLASD” (used for [5]) fail to inform the algorithm what is the active material in the detector. And, the key common issue leading to bias in 252 Cf related to neutron detectors is not whether it is a glass detector or a scintillator. No, the issue is that 6 Li was within both detector types and that even small mistakes in the detector response functions around approximately 200 keV are amplified by the 6 Li(n,α) resonance there leading to bias in data as highlighted in Fig. 1 and Ref. [1]. Hence, the features describing the neutron detector must call out the active material in the detector, rather than the existing EXFOR detector keyword, that the ML algorithm can find physically meaningful features related to bias. The AIACHNE team used a precursor of the WPEC (Working Party on International Nuclear Data Evaluation Co-operation) SG(Subgroup)-50 format to store the metadata for the ML analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Critical statistical assessment of data in metal additive manufacturing

Obtaining high quality data reflecting the relationships between the additive manufacturing (AM) process parameters, material microstructure and mechanical properties is crucial for the use of machine learning in AM. A database of over 4,000 data entries of metal AM was created thanks to a large number of literature studies on key process parameters and indicators of build quality. Meta-analysis reveals critical biases in the literature. Firstly, majority of studies report only high quality builds, these imbalances in reporting result in weak correlation between process parameters, properties and consolidation, limiting the ability of machine learning models to generalize beyond optimized conditions. Nevertheless, the trained models accurately predict yield strength ($R^2 = 0.85$), suggesting that certain process–property relationships are effectively captured within these models. Secondly, quantitative microstructural data are largely absent, limiting the learning of the microstructure-mechanical properties relationships. Finally, current process window identification is based largely on the consolidation, despite significant uncertainty in its measurement. It is important to identify the process map on the basis of not only the consolidation, but also mechanical behaviour under loading. Such a identification shows that 316 L and Inconel have much larger process map (i.e. highly printable) in comparison to the AlSi10Mg and Ti6Al4V.

Additive manufacturing↗

NEWTS Argonne Geothermal Geochemical Database with CoDART

A database of geochemical compositions of aqueous species in potential geothermal resources. The NETL NEWTS team has formatted the original Argonne V2 database for Charge Balance, Input into OLI Studio, and Input into GWB. (Geochemist WorkBench) In addition, some missing species in the original database were predicted using machine learning techniques within CoDart software, a public ML software developed by the Nation Energy Technology Laboratory. We have made the Input into CoDart and one example output from CoDart available in this dataset.

Aqueous Chemistry↗

Physics-Guided Continual Learning for Predicting Emerging Aqueous Organic Redox Flow Battery Material Performance

Aqueous organic redox flow batteries (AORFBs) have gained popularity in renewable energy storage due to their low cost, environmental friendliness and scalability. The rapid discovery of aqueous soluble organic (ASO) redox-active materials necessitates efficient machine learning surrogates for predicting battery performance. The physics-guided continual learning (PGCL) method proposed in this study can incrementally learn data from new ASO electrolytes while addressing catastrophic forgetting issues in conventional machine learning. Using a AORFB database with a thousand potential materials generated by a 780 $\text{cm}^2$ interdigitated cell model, PGCL incorporates AORFB physics to optimize the continual learning task formation and training strategies to retain previously learned battery material knowledge. Finally, the trained PGCL demonstrates its capability in assessing emerging ASO materials within the established parameter space when evaluated with the dihydroxyphenazine isomers.

25 ENERGY STORAGE↗

AI-powered exploration of molecular vibrations, phonons, and spectroscopy

The vibrational dynamics of molecules and solids play a critical role in defining material properties, particularly their thermal behaviors. However, theoretical calculations of these dynamics are often computationally intensive, while experimental approaches can be technically complex and resource-demanding. Recent advancements in data-driven artificial intelligence (AI) methodologies have substantially enhanced the efficiency of these studies. This review explores the latest progress in AI-driven methods for investigating atomic vibrations, emphasizing their role in accelerating computations and enabling rapid predictions of lattice dynamics, phonon behaviors, molecular dynamics, and vibrational spectra. Key developments are discussed, including advancements in databases, structural representations, machine-learning interatomic potentials, graph neural networks, and other emerging approaches. Compared to traditional techniques, AI methods exhibit transformative potential, dramatically improving the efficiency and scope of research in materials science. The review concludes by highlighting the promising future of AI-driven innovations in the study of atomic vibrations.

Han, Bowen [Oak Ridge National Laboratory (ORNL), ↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fuel Cell Inverter Dataset

This data set contains the three phase AC voltage, three phase AC current, DC voltage and DC current. These data sets were captured during fuel cell inverter operation in grid-connected dispatch, islanded load changes, transition from grid-connected mode to islanded mode and vice-versa.

25 ENERGY STORAGE↗

Forming a database to study reversed magnetic shear from the National Spherical Torus eXperiment using machine learning

Achieving a long-lived reversed magnetic shear (RMS) target plasma in the National Spherical Torus eXperiment Upgrade will require developing various sustainment scenarios. To help with the ongoing plasma control efforts, the development of a new analysis for the motional Stark effect (MSE) diagnostic using a machine learning algorithm, namely, MSE-ML, is described. MSE-ML will be used to identify patterns during RMS discharges, some of which suffer magnetohydrodynamic (MHD) events resulting in current redistribution and monotonic q-profiles. A database consisting of q and magnetic shear profiles is being constructed primarily based on the existing National Spherical Torus eXperiment data with equilibrium reconstructions constrained by the magnetic field pitch angle profile measured using the multi-channel MSE diagnostic. An unsupervised k-means clustering of the data is developed to study the RMS formation as a function of time. The initial clustering from the q-profiles shows significant differences in both amplitude and the duration of the RMS period. As a goal, the clustering results that detect and distinguish shots with substantial and sustained RMS are to be used as a preprocessing step in a supervised algorithm to identify the underlying conditions that lead to long-lasting improved confinement with RMS. Another aim of the MSE-ML study is to identify precursors of RMS-destroying MHD events in either derived data such as the q-profile or directly measured data such as the magnetic field pitch angle profile.

Uzun-Kaymak, I. U. (ORCID:0000000276251493)↗

Macroscopic trends of neoclassical tearing stability in high-field H-mode tokamak pilot plants

The neoclassical tearing mode (NTM) stability metric—minimum marginally stable island width $w$$^{*}_{m}$—was compared across 14651 inductive high-field tokamak pilot plant equilibria. Larger devices with reduced elongation and/or increased minor radius demonstrated an order-of-magnitude increase in $w$$^{*}_{m}$, primarily due to a reduction in bootstrap drive. This work is part of an ongoing effort to ensure passive NTM-stability in the ARC tokamak, in which the technology to achieve active tearing-suppression with localised electron cyclotron current drive does not yet exist. The equilibrium scenarios in the database were Monte Carlo generated and normalised to the same >400MW fusion power, minimum pressure scenario at a range of plasma currents, before tearing analysis using the modified Rutherford equation was applied for all resonant poloidal and toroidal m, n modes up to n = 4. Single-helicity toroidal Δ' calculations in resistive DCON set the minimum marginally stable island width, and a simple modal scaling proportional to –m 2 n –1 was identified for high-m Δ' values. The dominant correlates of $w$$^{*}_{m}$ and Δ' across the database were analysed using interpretable machine learning techniques.

NTM seeding↗