Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

he exponential growth of data centers—driven by artificial intelligence and cloud computing—is reshaping the U.S. energy landscape, presenting urgent challenges and transformative opportunities for data center developers, nuclear energy providers, and utility operators. As data centers are projected to consume up to 12% of U.S. electricity by 2028, stakeholders must address rapid deployment needs, grid congestion, and the demand for reliable, high-quality power. This presentation explores the multifaceted barriers to integrating data centers with nuclear and utility infrastructure, including land use constraints, public perception, regulatory complexity, and workforce alignment. It highlights the distinct priorities and operational cultures of each sector, and the friction that arises from misaligned planning horizons and risk tolerances. We examine collaborative strategies such as co-siting, hybrid power-purchase agreements, unified community engagement, and innovative financing models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Doppler Backscattering Data Analysis and Integrated Modeling with OMFIT

One Modeling Framework for Integrated Tasks (OMFIT) is a widely used software tool in the magnetic fusion research community. OMFIT provides magnetic fusion energy researchers with a framework for the development of special-purpose physics modules. This paper describes an OMFIT physics module pertaining to the Doppler Backscattering (DBS) fusion plasma diagnostic. DBS measures density fluctuations and flow velocity through plasma scattering of electromagnetic waves. The OMFIT DBS module was developed to analyze experimental DBS data and facilitate modeling of DBS systems installed on multiple tokamak devices. The OMFIT DBS module is designed to support several analysis workflows: detailed analysis of experimental data, experimental planning, and theory-based synthetic diagnostic modeling. The DBS module uses integrated modeling by leveraging other OMFIT physics modules to perform tasks related to DBS, e.g. ray/beam–tracing simulations, edge-localized mode–synchronized data analysis, magnetic equilibrium reconstruction, and fitting kinetic profile data. Furthermore, this paper describes several supported workflows and serves a reference for the OMFIT DBS module.

Doppler backscattering↗

Comparing optical four-flux model results with experimental data obtained by integrating sphere measurements

Four-flux theory is a way to model scattering through multiple layers of a system based on diffuse and collimated properties. When compared with measurement results obtained using an integrating sphere or a goniophotometer, an approximation is often made as the physical instrument cannot separate the collimated component from the diffuse light scattered in the forward direction. This paper tries to clarify the meaning of the word diffuse for the different cases and outlines simple corrections to improve the accuracy when comparing four-flux models and measured data based on sample haze and the geometry of the integrating sphere.

Bilokur, Maryna (ORCID:0000000191839650)↗

Energy Materials Chemistry Integrating Theory, Experiment and Data Science (Final Report)

The Energy Materials Chemistry Integrating Theory, Experiment and Data Science (EM-CITED) project is a multidisciplinary research effort focused on accelerating discovery of scientific knowledge via incorporation of data science and artificial intelligence in materials chemistry research. The project aims to advance materials chemistry-aware data science to unify theory and experiment knowledge streams. The work resulted in foundational AI frameworks for materials chemistry – Deep Reasoning Networks (DRNets), Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP), and Material-to-Spectrum (Mat2Spec) prediction – as well as a host of strategies for accelerated scientific discoveries through principled incorporation of data science in computational and experimental research.

36 MATERIALS SCIENCE↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Crossing the Finish Line: Integration of Data-Driven Process Control for Maximization of Energy and Resource Efficiency in Advanced Water Resource Recovery Facilities

Improvements in process monitoring and control at water resource recovery facilities (WRRFs) could result in reductions in electricity consumption, chemical inputs, and greenhouse gas emissions, as well as improved energy recovery. Many current WRRF data collection, monitoring, and control approaches use 20th century process monitoring and control systems, which require large design safety factors to ensure reliability in the absence of more advanced, precise controls. Implementation of more modern data-driven control tools could lead to more efficient operations that provide intrinsic reliability with better overall process performance at full-scale. This project (1) developed and demonstrated data-driven process controls at full-scale facilities for five promising WRRF process technologies that provide whole-plant approaches and offer substantial energy and resource recovery benefits, and (2) created a Machine Learning (ML) Toolkit and an implementation guide of new process control approaches that walks users through each step of the ML workflow and illustrates the steps through case study examples.

54 ENVIRONMENTAL SCIENCES↗

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING↗

Atmospheric Radiation Measurement (ARM) airborne field campaign data products between 2013 and 2018

Airborne measurements are pivotal for providing detailed, spatiotemporally resolved information about atmospheric parameters and aerosol and cloud properties, thereby enhancing our understanding of dynamic atmospheric processes. For 30 years, the US Department of Energy (DOE) Office of Science supported an instrumented Gulfstream 1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) Data Center and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated dataset was recently developed covering the final 6 years of G-1 operations (2013 to 2018, https://doi.org/10.5439/1999133; Mei and Gaustad, 2024). The integrated dataset includes data collected from 236 flights (766.4 h), which covered the Arctic, the US Southern Great Plains (SGP), the US West Coast, the eastern North Atlantic (ENA), the Amazon Basin in Brazil, and the Sierras de Córdoba range in Argentina. These comprehensive data streams provide much-needed insight into spatiotemporal variability in the thermodynamic quantities and aerosol and cloud properties for addressing essential science questions in Earth system process studies. This paper describes the DOE ARM merged G-1 datasets, including information on the acquisition, data collection challenges and future potentials, and quality control processes. It further illustrates the usage of this merged dataset to evaluate the Energy Exascale Earth System Model (E3SM) with the Earth System Model Aerosol–Cloud Diagnostics (ESMAC Diags) package.

54 ENVIRONMENTAL SCIENCES↗

Integrated Energy-Water Data for Cross-Sector Resilience

This white paper focuses on the “energy-for-water” domain, addressing the urgent need for integrated, empirical data to support regional management, benchmarking, and research on improving efficiency and developing technologies for water and wastewater management systems. The costs and energy required for the supply, treatment, and distribution of water and wastewater lack a standard data collection mechanism and centralized database or storage infrastructure, limiting data-driven decision-making across interdependent infrastructure systems.

42 ENGINEERING↗

A Feasibility Study on the Integration of Human Performance Data From Diverse Sources Based on the Complexity of a Proceduralized Task

Securing the safety of socio-technical systems including nuclear facilities is the upmost goal to ensure their sustainability because historical records demonstrate that the performance degradation of human operators (e.g., human errors) is one of the crucial contributors to the occurrence of unexpected events resulting in extensive casualties and financial losses. This implies that the collection of human performance data in diverse conditions with which they could be faced during the operation of nuclear facilities. As this collection requires significant resources, it is necessary to resolve how to accomplish it with limited resources. To address this challenge, as suggested in the SHEEP framework, it is indispensable to extract valuable insights after integrating various kinds of human performance data obtained from different sources. However, a practical method to soundly integrate them seems to be still incomplete. Accordingly, the applicability of TACOM (Task Complexity) measure is investigated as a tool to identify useful information based on the integration of human performance data observed from different simulation conditions. As a result, it is expected that the TACOM measure would play an important role in addressing the technical challenge in securing human performance data.

99 GENERAL AND MISCELLANEOUS↗

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigate whether integrating genomic, transcriptomic, and methylomic data can improve prediction for six Arabidopsis traits. We find that transcriptome- and methylome-based models have performances comparable to those of genome-based models. However, models built for flowering time using different omics data identify different benchmark genes. Nine additional genes identified as important for flowering time from our models are experimentally validated as regulating flowering. Gene contributions to flowering time prediction are accession-dependent and distinct genes contribute to trait prediction in different genotypes. Models integrating multi-omics data perform best and reveal known and additional gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

59 BASIC BIOLOGICAL SCIENCES↗

Integrated Framework of Multisource Data Fusion for Outage Location in Looped Distribution Systems

Accurate outage location is essential for expediting post-outage power restoration, minimizing outage duration, and enhancing the resilience of distribution networks. With the advent of advanced metering infrastructure, data-driven outage location methods have significantly advanced beyond traditional approaches that rely on manual inspections. However, existing methods still face critical challenges, like reliance on single-source data, limited ability to handle partially observable systems or difficulties with loop networks. To the best of our knowledge, no single approach has comprehensively addressed all of these challenges at once. To this end, this paper proposes a comprehensive multisource data fusion framework for outage locations via probabilistic graph networks. The framework consists of three key phases. First, a novel method for reconstituting distribution networks with loops is developed, transforming looped networks into multiple radial subnetworks that retain all outage causalities of the original network. Second, Bayesian network (BN) models are established for each subnetwork, integrating multiple data sources and network structures. Finally, a joint Gibbs sampling mechanism, featuring forward and backward information flow, is designed to merge data from separate BN models and maximize the utilization of limited evidence, ensuring accurate outage location identification. In conclusion, the framework was validated on two modified public test systems, and comparative studies confirmed its effectiveness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Qualification of Digitized Legacy Fast Reactor Data

The Integral Fast Reactor (IFR) fuel compatibility test program (1984-1994) included a variety of fuel pin examinations conducted at the Hot Fuel Examination Facility (HFEF) and the Alpha-Gamma Hot Cell Facility (AGHCF). Hard copy data records of these examinations have been recovered, scanned, and preserved in PDF format. Many hard copy records are now qualified in accordance with an NRC-approved Quality Assurance Program Plan (QAPP), and there is an ongoing effort to qualify additional legacy records. This legacy fuel performance data is vital to support design and licensing of fast reactors with validation of state-of-the-art codes and advanced methods for design and analysis. Stakeholders can most easily utilize this data when the PDF scans have been converted into digital data tables. However, qualification of the scanned hard copy data does not qualify the digital data file resulting from the digitization of the data contained in the record; the subject matter expert (SME) must make a review of the digitized data table as well before it can be designated as qualified. This report outlines a peer review process to qualify the digital data file(s), typically in CSV format, corresponding to hard copy records in accordance with the existing QAPP.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗