Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-centric machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Data-centric machine learning in quantum information science

Abstract We propose a series of data-centric heuristics for improving the performance of machine learning systems when applied to problems in quantum information science. In particular, we consider how systematic engineering of training sets can significantly enhance the accuracy of pre-trained neural networks used for quantum state reconstruction without altering the underlying architecture. We find that it is not always optimal to engineer training sets to exactly match the expected distribution of a target scenario, and instead, performance can be further improved by biasing the training set to be slightly more mixed than the target. This is due to the heterogeneity in the number of free variables required to describe states of different purity, and as a result, overall accuracy of the network improves when training sets of a fixed size focus on states with the least constrained free variables. For further clarity, we also include a ‘toy model’ demonstration of how spurious correlations can inadvertently enter synthetic data sets used for training, how the performance of systems trained with these correlations can degrade dramatically, and how the inclusion of even relatively few counterexamples can effectively remedy such problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

A data-centric weak supervised learning for highway traffic incident detection

Using the data from loop detector sensors for near-real-time detection of traffic incidents on highways is crucial to averting major traffic congestion. While recent supervised machine learning methods offer solutions to incident detection by leveraging human-labeled incident data, the false alarm rate is often too high to be used in practice. Specifically, the inconsistency in the human labeling of the incidents significantly affects the performance of supervised learning models. To that end, we focus on a data-centric approach to improve the accuracy and reduce the false alarm rate of traffic incident detection on highways. We develop a weak supervised learning workflow to generate high-quality training labels for the incident data without the ground truth labels, and we use those generated labels in the supervised learning setup for final detection. This approach comprises three stages. First, we introduce a data preprocessing and curation pipeline that processes traffic sensor data to generate high-quality training data through leveraging labeling functions, which can be domain knowledge-related or simple heuristic rules. Second, we evaluate the training data generated by weak supervision using three supervised learning models-random forest, k-nearest neighbors, and a support vector machine ensemble-and long short-term memory classifiers. The results show that the accuracy of all of the models improves significantly after using the training data generated by weak supervision. Third, we develop an online real-time incident detection approach that leverages the model ensemble and the uncertainty quantification while detecting incidents. Finally, we show that our proposed weak supervised learning workflow achieves a high incident detection rate (0.90) and low false alarm rate (0.08).

97 MATHEMATICS AND COMPUTING↗

Data-Centric Operational Design Domain Characterization for Machine Learning-Based Aeronautical Products

We give for Machine Learning (ML)-based aeronautical products, a first rigorous characterization of Operational Design Domains (ODDs). Unlike in other application sectors (such as self-driving road vehicles) where ODD development is scenario-based, our approach is data-centric: we propose the dimensions along which the parameters that define an ODD can be explicitly captured, using a top-down approach starting from system specifications, and a bottom-up approach starting from detailed ML Model (MLM) designs. Then we give a categorization of the data that ML-based applications can encounter in operation, identifying their system-level relevance and impact. Specifically, we discuss how those data categories are useful to determine: (1) the requirements necessary to drive the design of MLMs; (2) the potential effects on the MLM and higher levels of the system hierarchy; (3) the learning assurance processes that may be needed, and (4) system architectural considerations. We illustrate the underlying concepts with an example of an aircraft flight envelope. The approach in this paper is one of the cornerstones of a future process guidance for development and certification/approval of safety-related aeronautical products implementing Artificial Intelligence (AI), currently being developed through aviation industry-based consensus, jointly by the SAE G-34 Committee for AI in aviation, and EUROCAE WG-114 for AI.

Aeronautical products↗

Data-centric framework for crystal structure identification in atomistic simulations using machine learning

Atomic-level modeling performed at large scales enables the investigation of mesoscale materials properties with atom-by-atom resolution. The spatial complexity of such cross-scale simulations renders them unsuitable for simple human visual inspection. Instead, specialized structure characterization techniques are required to aid interpretation. These have historically been challenging to construct, requiring significant intuition and effort. Here we propose an alternative framework for a fundamental structural characterization task: classifying atoms according to the crystal structure to which they belong. Our approach is data-centric and favors the employment of Machine Learning over heuristic rules of classification. A group of data-science tools and simple local descriptors of atomic structure are employed together with an efficient synthetic training set. We also introduce the first standard and publicly available benchmark data set for evaluation of algorithms for crystal-structure classification. Further, it is demonstrated that our data-centric framework outperforms all of the most popular heuristic methods—especially at high temperatures when lattices are the most distorted—while introducing a systematic route for generalization to new crystal structures. Moreover, through the use of outlier detection algorithms our approach is capable of discerning between amorphous atomic motifs (i.e., noncrystalline phases) and unknown crystal structures, making it uniquely suited for exploratory materials synthesis simulations.

36 MATERIALS SCIENCE↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

Data-Centric AI and the Open Energy Data Initiative (OEDI)

This presentation emphasizes the critical importance of data-centric AI. The limitations of model-centric AI when dealing with poor or insufficient data are highlighted, and it is illustrated how training models on inaccurate or noisy data leads to suboptimal results. This talk advocates for a hybrid approach that combines a focus on data quality and model parameters to achieve optimal results. The Open Energy Data Initiative (OEDI) is introduced as a valuable resource for obtaining high-quality energy-related datasets, hosting nearly 2,000 publicly accessible datasets, including 99 solar-related datasets, totaling over 2.7 petabytes of data. OEDI's data lakes enable users to query and work with data without extensive transfers. In conclusion, the significance of data-centric AI and adherence to data curation best practices is emphasized, positioning OEDI as a prime source of high-quality data for AI and machine learning in the renewable energy sector.

AI↗

Realizing the data-driven, computational discovery of metal-organic framework catalysts

Metal-organic frameworks (MOFs) have been widely investigated for challenging catalytic transformations due to their well-defined structures and high degree of synthetic tunability. These features, at least in principle, make MOFs ideally suited for a computational approach towards catalyst design and discovery. Nonetheless, the widespread use of data science and machine learning to accelerate the discovery of MOF catalysts has yet to be substantially realized. In this review, we provide an overview of recent work that sets the stage for future high-throughput computational screening and machine learning studies involving MOF catalysts. This is followed by a discussion of several challenges currently facing the broad adoption of data-centric approaches in MOF computational catalysis, and we share possible solutions that can help propel the field forward.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Operational Technology Behavioral Analytics (OTBA) (Final Technical Report DE-FE0031640)

This final report provides a summary of the methodology, findings, lessons learned, and insights from an investigation into the feasibility of the Operational Technology Behavioral Analytics (OTBA) cybersecurity approach. The concept was evaluated with data from the National Carbon Capture Center (NCCC) – a U.S. Department of Energy (DOE) funded facility that is managed and operated by Southern Company Services, Inc. at Alabama Power Company’s E. C. Gaston generating power plant in Wilsonville, Alabama. Appropriate data sources for the post-combustion carbon capture system were identified. Infrastructure was deployed to monitor, capture and archive data for the system. Critical parameters for each subsystem were identified and analyzed. Machine-learning algorithms were used to establish and characterize normal operations and subsequently identify anomalies. This effort yielded valuable insights and formed the basis of a data-centric strategy for detecting cyber-attacks along with a coordinated response philosophy. A significant takeaway is that the OTBA cybersecurity approach is quite portable; it can be applied to other critical infrastructure beyond fossil power generation.

20 FOSSIL-FUELED POWER PLANTS↗

OPERATIONAL TECHNOLOGY BEHAVIORAL ANALYTICS (OTBA) – A DATA-CENTRIC APPROACH FOR REDUCING CYBERSECURITY RISK

This paper provides a summary of the methodology, findings, lessons learned, and insights from an investigation into the feasibility of the Operational Technology Behavioral Analytics (OTBA) cybersecurity approach. The concept was evaluated with data from the National Carbon Capture Center (NCCC) – a U.S. Department of Energy (DOE) funded facility that is managed and operated by Southern Company at Alabama Power’s E. C. Gaston generating power plant in Wilsonville, Alabama. Appropriate data sources for the post-combustion carbon capture system were identified. Infrastructure was deployed to monitor, capture and archive data for the system. Critical parameters for each subsystem were identified and analyzed. Machine-learning algorithms were used to establish and characterize normal operations and subsequently identify anomalies. This effort yielded valuable insights and formed the basis of a data-centric strategy for detecting cyber-attacks along with a coordinated response philosophy. A significant takeaway is that the OTBA cybersecurity approach is quite portable; it can be applied to other critical infrastructure beyond fossil power generation.

Black, Clifton↗

Learning and Controlling Silicon Dopant Transitions in Graphene Using Scanning Transmission Electron Microscopy

A machine learning approach is introduced to determine the transition dynamics of silicon atoms on a single layer of carbon atoms, when stimulated by the electron beam of a scanning transmission electron microscope (STEM). This method is data-centric, leveraging data collected on a STEM. The data samples are processed and filtered to produce symbolic representations, which is used to train a neural network to predict transition probabilities. These learned transition dynamics are then leveraged to guide a single silicon atom throughout the lattice to pre-determined target destinations. Empirical analyses are presented that demonstrate the efficacy and generality of the approach.

36 MATERIALS SCIENCE↗

Lockdown impacts on residential electricity demand in India: A data-driven and non-intrusive load monitoring study using Gaussian mixture models

This study evaluates the effect of complete nationwide lockdown in 2020 on residential electricity demand across 13 Indian cities and the role of digitalisation using a public smart meter dataset. We undertake a data-driven approach to explore the energy impacts of work-from-home norms across five dwelling typologies. Our methodology includes climate correction, dimensionality reduction and machine learning-based clustering using Gaussian Mixture Models of daily load curves. Results show that during the lockdown, maximum daily peak demand increased by 150-200% as compared to 2018 and 2019 levels for one room-units (RM1), one bedroom-units (BR1) and two bedroom-units (BR2) which are typical for low- and middle-income families. While the upper-middle- and higher-income dwelling units (i.e., three (3BR) and more-than-three bedroom-units (M3BR)) saw night-time demand rise by almost 44% in 2020, as compared to 2018 and 2019 levels. Our results also showed that new peak demand emerged for the lockdown period for RM1, BR1 and BR2 dwelling typologies. We found that the lack of supporting socioeconomic and climatic data can restrict a comprehensive analysis of demand shocks using similar public datasets, which informed policy implications for India's digitalisation. We further emphasised improving the data quality and reliability for effective data-centric policymaking.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data-Centric Approach to Capture Non-Polynomial Nonlinear Dynamics

We propose an analytical construction of observable functions in the extended dynamic mode decomposition (EDMD) algorithm. EDMD is a numerical method for approximating the spectral properties of the Koopman operator. The choice of observable functions is fundamental for applying EDMD to nonlinear problems arising in systems and control. Existing methods either start from a set of dictionary functions and look for the subset that best fits the underlying nonlinear dynamics or rely on machine learning algorithms to “learn” observable functions. Conversely, in this paper, we start from the dynamical system model and lift it through the Lie derivatives, rendering it into a polynomial form. This proposed transformation into a polynomial form is exact and provides an adequate set of observable functions. The strength of the proposed approach is its applicability to a broader class of nonlinear dynamical systems, particularly those with nonpolynomial functions and compositions thereof. Moreover, it retains the physical interpretability of the underlying dynamical system and can be readily integrated into existing numerical libraries. We demonstrate the proposed approach with an application to electric power systems. The modeled system consists of a single generator connected to an infinite bus, where nonlinear terms include sine and cosine functions. The results demonstrate the effectiveness of the proposed procedure in off-attractor nonlinear dynamics for estimation and prediction; the observable functions obtained from the proposed construction outperform methods that use dictionary functions comprising monomials or radial basis functions.

extended dynamic mode decomposition↗

Understanding and design of interstitial oxygen conductors

Highly efficient oxygen-active materials that react with, absorb, and transport oxygen is essential for fuel cells, electrolyzers and related applications. While vacancy-mediated oxygen-ion conductors have long been the focus of research, they are limited by high migration barriers at intermediate temperatures (400–600 °C), which hinder their practical applications. In contrast, interstitial oxygen conductors exhibit significantly lower migration barriers enabling higher ionic conductivity at lower temperatures. This review systematically examines both well-established and recently identified families of interstitial oxygen-ion conductors, focusing on how their unique structural motifs such as corner-sharing polyhedral frameworks, isolated polyhedral, and cage-like architectures, facilitate low migration barriers through interstitial and/or interstitialcy diffusion mechanisms. A central discussion of this review focuses on the evolution of design strategies, from targeted donor doping, element screening, to physical-intuition descriptor material screening and machine learning approach, which leverage computational tools to explore vast chemical spaces in search for new interstitial conductors. The success of these strategies demonstrates that a significant, largely unexplored space remains for discovering high-performing interstitial oxygen conductors. Crucial features enabling high-performance interstitial oxygen diffusion include the availability of electrons for oxygen reduction and sufficient structural flexibility with accessible volume for interstitial accommodation and migration. This review concludes with a forward-looking perspective, proposing a knowledge-driven methodology that integrates current understanding with data-centric approaches to identify promising interstitial oxygen conductors outside traditional search paradigms. These approaches are expected to significantly accelerate the development of high-performance interstitial oxygen conductors for a variety of oxygen-active applications, ultimately paving the way for more efficient and sustainable energy technologies.

Interstitial oxygen conductors↗

Automation-Accelerated Electrolyte Design Mitigates Solubility Competition between Redox-Active Molecules and Supporting Salts

In nonaqueous redox-flow batteries (NRFBs), redox-active organic molecules (ROMs) and supporting salts compete for solvation sites, limiting achievable energy density. We combine automated high-throughput experimentation (HTE) with camera-based saturation monitoring and quantitative NMR to measure paired (ROM, salt) solubilities across single and mixed organic solvents. Using 2,1,3-benzothiadiazole (BTZ) with lithium bis(trifluoromethanesulfonyl)imide (LiTFSI) as a model system, we find that a binary m-xylene/acetonitrile mixture dissolves ≈3 M of both BTZ and LiTFSI─surpassing the previously reported 2 M ceiling for neat acetonitrile─by leveraging complementary solvation (MX is BTZ-philic and salt-phobic; ACN stabilizes LiTFSI). A random-forest model (RMSE ≈ 0.24) trained on solvent descriptors highlights log P and salt concentration as dominant predictors and predicts MX/ACN ≈0.3/0.7 (v/v) to be near-optimal. These formulations retain practical viscosity and ∼5 mS·cm –1 conductivity at high loading. In conclusion, the workflow provides a reproducible, data-centric route to NRFB electrolyte design and motivates an open, standardized dual-solute solubility resource for accelerated electrolyte discovery.

Electrolytes↗

Machine learning in nuclear materials research

Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes with associated transmutations, high temperature and temperature gradients, mechanical stresses, and corrosive coolants. They also have a wide range of microstructural and chemical makeups, resulting in multifaceted and often out-of-equilibrium interactions. Machine learning (ML) is increasingly being used to tackle these complex time-dependent interactions and aid researchers in developing models and making predictions, sometimes with better accuracy than traditional modeling that focuses on one or two parameters at a time. Conventional practices of acquiring new experimental data in nuclear materials research are often slow and expensive, limiting the opportunity for data-centric ML, but new methods are changing that paradigm. Here we review high-throughput computational and experimental data approaches, especially robotic experimentation and active learning that is based on Gaussian process and Bayesian optimization. We show ML examples in structural materials (e.g., reactor pressure vessel (RPV) alloys and radiation detecting scintillating materials) and highlight new techniques of high-throughput sample preparation and characterizations, and automated radiation/environmental exposures and real-time online diagnostics. Herein, this review suggests that ML models of material constitutive relations in plasticity, damage, and even electronic and optical responses to radiation are likely to become powerful tools as they develop. Finally, we speculate on how the recent trends of using natural language processing (NLP) to aid the collection and analysis of literature data, interpretable artificial intelligence (AI), and the use of streamlined scripting, database, workflow management, and cloud computing platforms that will soon make the utilization of ML techniques as commonplace as the spreadsheet curve-fitting practices of today.

36 MATERIALS SCIENCE↗

MICCO: An Enhanced Multi-GPU Scheduling Framework for Many-Body Correlation Functions

Calculation of many-body correlation functions is one of the critical kernels utilized in many scientific computing areas, especially in Lattice Quantum Chromodynamics (Lattice QCD). It is formalized as a sum of a large number of contraction terms each of which can be represented by a graph consisting of vertices describing quarks inside a hadron node and edges designating quark propagations at specific time intervals. Due to its computation- and memory-intensive nature, real-world physics systems (e.g., multi-meson or multi-baryon systems) explored by Lattice QCD prefer to leverage multi-GPUs. Different from general graph processing, many-body correlation function calculations show two specific features: a large number of computation-/data-intensive kernels and frequently repeated appearances of original and intermediate data. The former results in expensive memory operations such as tensor movements and evictions. The latter offers data reuse opportunities to mitigate the data-intensive nature of many-body correlation function calculations. However, existing graph-based multi-GPU schedulers cannot capture these data-centric features, thus resulting in a sub-optimal performance for many-body correlation function calculations. To address this issue, this paper presents a multi-GPU scheduling framework, MICCO, to accelerate contractions for correlation functions particularly by taking the data dimension (e.g., data reuse and data eviction) into account. This work first performs a comprehensive study on the interplay of data reuse and load balance, and designs two new concepts: local reuse pattern and reuse bound to study the opportunity of achieving the optimal trade-off between them. Based on this study, MICCO proposes a heuristic scheduling algorithm and a machine-learning-based regression model to generate the optimal setting of reuse bounds. Specifically, MICCO is integrated into a real-world Lattice QCD system, Redstar, for the first time running on multiple GPUs. The evaluation demonstrates MICCO outperforms other state-of-art works, achieving up to 2.25× speedup in synthesized datasets, and 1.49× speedup in real-world correlation functions.

Wang, Qihan↗

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.

cloud-optimized↗