Search NASA⌕ Search

SEARCH · Search NASA

Results for “data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Gray-Box Modeling for Distribution Systems With Inverter-Based Resources: Integrating Physics-Based and Data-Driven Approaches

Here, in this paper, we develop a novel gray-box modeling approach for distribution systems with inverter-based resources (IBRs). The proposed gray-box modeling method aims to improve estimation accuracy by taking advantages of both physics-based (white-box) and data-driven (black-box) modeling approaches. To this end, we utilize partial physical knowledge of the system, including the inverters’ structures and control diagrams, as well as the equivalent network model simplified through Kron reduction. The white-box model containing unknown parameters is then constructed with mathematical equations and an optimization-based method is subsequently employed to identify these unknown parameters within the white-box model. Next, the graybox modeling framework is then constructed by embedding the output variables of the white-box model into the input vector of a black-box model (represented using a neural network). Finally, the black-box section is trained using the collected input-output datasets and the gray-box model is then obtained. Furthermore, case studies demonstrate that our gray-box modeling approach effectively improves estimation accuracy compared to purely physics-based or data-driven methods.

42 ENGINEERING↗

2015-2017 California Vehicle Survey

The 2015-2017 California Vehicle Survey of residential and commercial light-duty vehicle owners in California assessed consumer preferences for vehicles and included a targeted sample of plug-in electric vehicle (PEV) owners. Resource Systems Group conducted the survey on behalf of the California Energy Commission. In addition to economic and demographic data, the survey integrated light-duty vehicle holding and use information with vehicle choice data collected via the stated preferences survey's set of eight vehicle and fuel type choice exercises. The PEV owner survey participants provided additional data on charging behavior, electricity rates, and their main motivations for purchasing PEVs.

1Hz data↗

2019 California Vehicle Survey

The 2019 California Vehicle Survey of residential and commercial light-duty fleet owners in California assessed consumer preferences for vehicles and included a targeted sample of plug-in electric vehicle (PEV) owners. In addition to economic and demographic data, the survey integrated light-duty vehicle holding and use information with vehicle choice data, which was collected via a set of eight exercises on vehicle and fuel type choice. The PEV owner survey participants provided additional data on charging behavior, electricity rates, and their main motivations for purchasing PEVs.

1Hz data↗

Integration of Condition-Based, Diagnostic, Prognostic, And Anomaly Detection Data into Reliability Models to Support a Predictive Maintenance Context

Reliability data employed in plant reliability models are an approximated integral representation of the past industrywide operational experience, and they neglect the present asset health status (available, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Ideally, in a predictive maintenance context, system reliability models should support decision making by propagating actual health information from the asset to the system level in order to provide a quantitative snapshot of system health and identify the most critical assets. Asset health should be informed solely by that specific asset’s current and historical performance data and should not be an approximated integral representation of the past industrywide operational experience (as currently performed by system reliability models through Bayesian updating processes). This paper proposes a reliability modeling approach that relies on asset diagnostic and prognostic assessments, along with monitoring data to measure asset health. We show how state-of-the art condition-based, diagnostic, prognostic, and anomaly detection models can be linked to system reliability models not in probability terms, but in terms of margin where margin is defined as the “distance” between the present status and an undesired event (e.g., failure or unacceptable performance). Then, we show how the propagation of margin data from the asset to the system level is performed through classical reliability models such as fault trees or reliability block diagrams. The described method is in fact able to propagate heterogenous health data from the asset to the system level in order to analytically assess system health.

97 MATHEMATICS AND COMPUTING↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗

Next-Generation Materials Design: Quantum Mechanics and Data-Driven Modeling

The future of materials design is rapidly advancing through the combination of quantum mechanics and data-driven modeling. These approaches integrate quantum principles with advanced data analysis, enabling precise insights into material behavior. This talk will highlight recent progress in using these methods for computational design, particularly in high-entropy alloy catalysts, emphasizing the role of hierarchical machine-learning architectures for accurate predictions. Additionally, I will discuss our work on developing machine learning interatomic potentials (MLPs) for single-element metals, metal oxides, and alloys under extreme conditions, focusing on melting behavior and phase properties at high temperatures and pressures. We have also refined our MLP models to capture dynamic surface interactions, such as CO2 and CO adsorption on MgO, using both static and molecular dynamics simulations. These models maintain high accuracy while significantly reducing computational costs compared to first-principles calculations. By enabling efficient and accurate simulations, this work supports broader community adoption, optimizes datasets for materials discovery, and extends the accessible time, size, and environmental conditions beyond the limits of experiments and traditional simulations.

machine learning↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Integration of Solid Oxide Fuel Cell Systems Into Artificial Intelligence Data Centers

This report presents the results of a techno-economic analysis (TEA) that evaluates the economic benefits of integrating solid oxide fuel cell (SOFC) systems with artificial intelligence (AI) data centers. The analysis was completed in two phases: a scoping-level analysis was performed to identify impactful integration opportunities, followed by a more detailed TEA. Results show that, due to their modularity, SOFC can meet the 99.999% availability requirement of data centers with minimal additional costs. Heat integration via absorption chillers decreases data center electricity consumption at the tradeoff of increased water consumption. Higher SOFC exhaust temperatures are important for achieving larger electricity savings. Finally, power electronics integration with SOFC direct current electricity can reduce electricity consumption by 9 percent and reduce water consumption by 6.4 percent.

20 FOSSIL-FUELED POWER PLANTS↗

Illuminating the Night: A Survey of Super-Resolution Methods for Nighttime Light Images

Nighttime Light (NTL) images provide critical insights into urbanization, disaster response, and energy consumption. The VIIRS Day/Night Band (DNB) sensor offers high-quality NTL imagery with daily revisit rates, but the available spatial resolution hinders fine-grained accurate analysis. Super-resolution techniques aim to increase the resolution of NTL images, enabling more detailed assessments of infrastructure, light pollution, economic activity, and power outages. However, existing state-of-the-art super-resolution methods designed for natural images struggle with the unique characteristics of NTL data. This work provides a comprehensive review of super-resolution methods across multiple image modalities, evaluates their effectiveness on VI-IRS DNB data, and proposes a multi-modal super-resolution approach tailored to NTL imagery. The proposed approach integrates VIIRS DNB data with road networks and land use information to improve reconstruction accuracy and spatial detail. Code is available for this project athttps://code.ornl.gov/viirs-sr/sr-demos.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Illuminating the Night: A Survey of Super-Resolution Methods for Nighttime Light Images

Nighttime Light (NTL) images provide critical insights into urbanization, disaster response, and energy consumption. The VIIRS Day/Night Band (DNB) sensor offers high-quality NTL imagery with daily revisit rates, but the available spatial resolution hinders fine-grained accurate analysis. Super-resolution techniques aim to increase the resolution of NTL images, enabling more detailed assessments of infrastructure, light pollution, economic activity, and power outages. However, existing state-of-the-art super-resolution methods designed for natural images struggle with the unique characteristics of NTL data. This work provides a comprehensive review of super-resolution methods across multiple image modalities, evaluates their effectiveness on VIIRS DNB data, and proposes a multi-modal super-resolution approach tailored to NTL imagery. The proposed approach integrates VIIRS DNB data with road networks and land use information to improve reconstruction accuracy and spatial detail. Code is available for this project at https://code.ornl.gov/viirs-sr/sr-demos.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Scalable Real-Time Data Assimilation Framework for Predicting Turbulent Atmosphere Dynamics

AI-based foundation models like FourCastNet, GraphCast are revolutionizing weather and climate predictions but are not yet ready for operational use. Their limitation lies in the absence of a data assimilation system to incorporate real-time Earth system observations, crucial for accurately forecasting events like tropical cyclones. To overcome these obstacles, we introduce a generic real-time data assimilation framework and demonstrate its end-to-end performance on the Frontier supercomputer. This framework comprises two primary modules: an ensemble score filter (EnSF), which significantly outperforms the state-of-the-art data assimilation method, and a vision transformer-based surrogate capable of real-time adaptation through the integration of observational data. We demonstrate both the strong and weak scaling of our framework up to 1024 GPUs on the Exascale supercomputer, Frontier. Our results not only illustrate the framework's exceptional scalability on high-performance computing systems, but also demonstrate the importance of supercomputers in real-time data assimilation for weather and climate predictions.

Lu, Dan↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Pipeline for Integrated Projects in Energy Systems (PIPES): A Tool for Integrated System Planning [Slides]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. This presentation introduces PIPES a multi-model tool for integrated system planning; it describes the underlying architecture, deep dives into common user workflows, and outlines the upcoming development roadmap beyond its current alpha state.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗