Search NASA⌕ Search

SEARCH · Search NASA

Results for “data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Integration of GOES Data for Solar Resource Assessment of the Contiguous United States

The National Solar Radiation Database (NSRDB), produced by the National Laboratory of the Rockies (NLR), provides high-resolution solar resource data for the contiguous United States (CONUS) using Geostationary Operational Environmental Satellite (GOES) East and West observations. This study evaluates the integration of multi-satellite data within the GOES-East/West overlap regions, where conventional longitude-based selection methods often produce an artificial boundary seam. Our results demonstrate that an advanced blending algorithm, which incorporates sun-satellite scattering angles and satellite viewing zenith angles, improves NSRDB accuracy and creates a spatially continuous dataset. Validation against ground-based irradiance measurements reveals reductions in both percentage error (PE) and normalized Root Mean Square Error (nRMSE), particularly in the central United States. The dynamical integration of multi-satellite data provides a robust foundation for more precise modeling of solar resource and improved spatiotemporal analysis of solar ramp across the CONUS.

14 SOLAR ENERGY↗

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Validating automated resonance evaluation with synthetic data

The integrity and precision of nuclear data are crucial for a broad spectrum of applications, from national security and nuclear reactor design to medical diagnostics, where the associated uncertainties can significantly impact outcomes. A substantial portion of uncertainty in nuclear data originates from the subjective biases in the evaluation process, a crucial phase in the nuclear data production pipeline. Recent advancements indicate that automation of certain routines can mitigate these biases, thereby standardizing the evaluation process and enhancing reproducibility. This research aims to provide a methodology, framework, and metrics for the validation of automated nuclear data evaluation software leveraging high-quality synthetic data that closely mimic real experimental observables. An introduced error metric provides a scale and intuitive measure of the evaluation quality by quantifying the estimate’s accuracy and performance across the specified energy range. Synthetic data provides access to experimental observables and underlying resonance parameters, enabling comparison of different evaluations. The methodology is demonstrated using Ta-181 isotope data in the resolved resonance region. The Automated Resonance Identification Subroutine (ARIS), which operates without prior resonance information, was used to test and showcase the framework’s capabilities utilizing the proposed error metrics. The results demonstrate the effectiveness of the proposed approach and framework for optimizing software parameters and testing hypotheses through “what-if” controlled experiments, such as modifying assumptions about experimental conditions or average resonance parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

he exponential growth of data centers—driven by artificial intelligence and cloud computing—is reshaping the U.S. energy landscape, presenting urgent challenges and transformative opportunities for data center developers, nuclear energy providers, and utility operators. As data centers are projected to consume up to 12% of U.S. electricity by 2028, stakeholders must address rapid deployment needs, grid congestion, and the demand for reliable, high-quality power. This presentation explores the multifaceted barriers to integrating data centers with nuclear and utility infrastructure, including land use constraints, public perception, regulatory complexity, and workforce alignment. It highlights the distinct priorities and operational cultures of each sector, and the friction that arises from misaligned planning horizons and risk tolerances. We examine collaborative strategies such as co-siting, hybrid power-purchase agreements, unified community engagement, and innovative financing models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Doppler Backscattering Data Analysis and Integrated Modeling with OMFIT

One Modeling Framework for Integrated Tasks (OMFIT) is a widely used software tool in the magnetic fusion research community. OMFIT provides magnetic fusion energy researchers with a framework for the development of special-purpose physics modules. This paper describes an OMFIT physics module pertaining to the Doppler Backscattering (DBS) fusion plasma diagnostic. DBS measures density fluctuations and flow velocity through plasma scattering of electromagnetic waves. The OMFIT DBS module was developed to analyze experimental DBS data and facilitate modeling of DBS systems installed on multiple tokamak devices. The OMFIT DBS module is designed to support several analysis workflows: detailed analysis of experimental data, experimental planning, and theory-based synthetic diagnostic modeling. The DBS module uses integrated modeling by leveraging other OMFIT physics modules to perform tasks related to DBS, e.g. ray/beam–tracing simulations, edge-localized mode–synchronized data analysis, magnetic equilibrium reconstruction, and fitting kinetic profile data. Furthermore, this paper describes several supported workflows and serves a reference for the OMFIT DBS module.

Doppler backscattering↗

Comparing optical four-flux model results with experimental data obtained by integrating sphere measurements

Four-flux theory is a way to model scattering through multiple layers of a system based on diffuse and collimated properties. When compared with measurement results obtained using an integrating sphere or a goniophotometer, an approximation is often made as the physical instrument cannot separate the collimated component from the diffuse light scattered in the forward direction. This paper tries to clarify the meaning of the word diffuse for the different cases and outlines simple corrections to improve the accuracy when comparing four-flux models and measured data based on sample haze and the geometry of the integrating sphere.

Bilokur, Maryna (ORCID:0000000191839650)↗

Energy Materials Chemistry Integrating Theory, Experiment and Data Science (Final Report)

The Energy Materials Chemistry Integrating Theory, Experiment and Data Science (EM-CITED) project is a multidisciplinary research effort focused on accelerating discovery of scientific knowledge via incorporation of data science and artificial intelligence in materials chemistry research. The project aims to advance materials chemistry-aware data science to unify theory and experiment knowledge streams. The work resulted in foundational AI frameworks for materials chemistry – Deep Reasoning Networks (DRNets), Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP), and Material-to-Spectrum (Mat2Spec) prediction – as well as a host of strategies for accelerated scientific discoveries through principled incorporation of data science in computational and experimental research.

36 MATERIALS SCIENCE↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Crossing the Finish Line: Integration of Data-Driven Process Control for Maximization of Energy and Resource Efficiency in Advanced Water Resource Recovery Facilities

Improvements in process monitoring and control at water resource recovery facilities (WRRFs) could result in reductions in electricity consumption, chemical inputs, and greenhouse gas emissions, as well as improved energy recovery. Many current WRRF data collection, monitoring, and control approaches use 20th century process monitoring and control systems, which require large design safety factors to ensure reliability in the absence of more advanced, precise controls. Implementation of more modern data-driven control tools could lead to more efficient operations that provide intrinsic reliability with better overall process performance at full-scale. This project (1) developed and demonstrated data-driven process controls at full-scale facilities for five promising WRRF process technologies that provide whole-plant approaches and offer substantial energy and resource recovery benefits, and (2) created a Machine Learning (ML) Toolkit and an implementation guide of new process control approaches that walks users through each step of the ML workflow and illustrates the steps through case study examples.

54 ENVIRONMENTAL SCIENCES↗

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING↗

Atmospheric Radiation Measurement (ARM) airborne field campaign data products between 2013 and 2018

Airborne measurements are pivotal for providing detailed, spatiotemporally resolved information about atmospheric parameters and aerosol and cloud properties, thereby enhancing our understanding of dynamic atmospheric processes. For 30 years, the US Department of Energy (DOE) Office of Science supported an instrumented Gulfstream 1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) Data Center and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated dataset was recently developed covering the final 6 years of G-1 operations (2013 to 2018, https://doi.org/10.5439/1999133; Mei and Gaustad, 2024). The integrated dataset includes data collected from 236 flights (766.4 h), which covered the Arctic, the US Southern Great Plains (SGP), the US West Coast, the eastern North Atlantic (ENA), the Amazon Basin in Brazil, and the Sierras de Córdoba range in Argentina. These comprehensive data streams provide much-needed insight into spatiotemporal variability in the thermodynamic quantities and aerosol and cloud properties for addressing essential science questions in Earth system process studies. This paper describes the DOE ARM merged G-1 datasets, including information on the acquisition, data collection challenges and future potentials, and quality control processes. It further illustrates the usage of this merged dataset to evaluate the Energy Exascale Earth System Model (E3SM) with the Earth System Model Aerosol–Cloud Diagnostics (ESMAC Diags) package.

54 ENVIRONMENTAL SCIENCES↗

Integrated Energy-Water Data for Cross-Sector Resilience

This white paper focuses on the “energy-for-water” domain, addressing the urgent need for integrated, empirical data to support regional management, benchmarking, and research on improving efficiency and developing technologies for water and wastewater management systems. The costs and energy required for the supply, treatment, and distribution of water and wastewater lack a standard data collection mechanism and centralized database or storage infrastructure, limiting data-driven decision-making across interdependent infrastructure systems.

42 ENGINEERING↗

A Feasibility Study on the Integration of Human Performance Data From Diverse Sources Based on the Complexity of a Proceduralized Task

Securing the safety of socio-technical systems including nuclear facilities is the upmost goal to ensure their sustainability because historical records demonstrate that the performance degradation of human operators (e.g., human errors) is one of the crucial contributors to the occurrence of unexpected events resulting in extensive casualties and financial losses. This implies that the collection of human performance data in diverse conditions with which they could be faced during the operation of nuclear facilities. As this collection requires significant resources, it is necessary to resolve how to accomplish it with limited resources. To address this challenge, as suggested in the SHEEP framework, it is indispensable to extract valuable insights after integrating various kinds of human performance data obtained from different sources. However, a practical method to soundly integrate them seems to be still incomplete. Accordingly, the applicability of TACOM (Task Complexity) measure is investigated as a tool to identify useful information based on the integration of human performance data observed from different simulation conditions. As a result, it is expected that the TACOM measure would play an important role in addressing the technical challenge in securing human performance data.

99 GENERAL AND MISCELLANEOUS↗