Search NASA⌕ Search

SEARCH · Search NASA

Results for “data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A Feasibility Study on the Integration of Human Performance Data From Diverse Sources Based on the Complexity of a Proceduralized Task

Securing the safety of socio-technical systems including nuclear facilities is the upmost goal to ensure their sustainability because historical records demonstrate that the performance degradation of human operators (e.g., human errors) is one of the crucial contributors to the occurrence of unexpected events resulting in extensive casualties and financial losses. This implies that the collection of human performance data in diverse conditions with which they could be faced during the operation of nuclear facilities. As this collection requires significant resources, it is necessary to resolve how to accomplish it with limited resources. To address this challenge, as suggested in the SHEEP framework, it is indispensable to extract valuable insights after integrating various kinds of human performance data obtained from different sources. However, a practical method to soundly integrate them seems to be still incomplete. Accordingly, the applicability of TACOM (Task Complexity) measure is investigated as a tool to identify useful information based on the integration of human performance data observed from different simulation conditions. As a result, it is expected that the TACOM measure would play an important role in addressing the technical challenge in securing human performance data.

99 GENERAL AND MISCELLANEOUS↗

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigate whether integrating genomic, transcriptomic, and methylomic data can improve prediction for six Arabidopsis traits. We find that transcriptome- and methylome-based models have performances comparable to those of genome-based models. However, models built for flowering time using different omics data identify different benchmark genes. Nine additional genes identified as important for flowering time from our models are experimentally validated as regulating flowering. Gene contributions to flowering time prediction are accession-dependent and distinct genes contribute to trait prediction in different genotypes. Models integrating multi-omics data perform best and reveal known and additional gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

59 BASIC BIOLOGICAL SCIENCES↗

Integrated Framework of Multisource Data Fusion for Outage Location in Looped Distribution Systems

Accurate outage location is essential for expediting post-outage power restoration, minimizing outage duration, and enhancing the resilience of distribution networks. With the advent of advanced metering infrastructure, data-driven outage location methods have significantly advanced beyond traditional approaches that rely on manual inspections. However, existing methods still face critical challenges, like reliance on single-source data, limited ability to handle partially observable systems or difficulties with loop networks. To the best of our knowledge, no single approach has comprehensively addressed all of these challenges at once. To this end, this paper proposes a comprehensive multisource data fusion framework for outage locations via probabilistic graph networks. The framework consists of three key phases. First, a novel method for reconstituting distribution networks with loops is developed, transforming looped networks into multiple radial subnetworks that retain all outage causalities of the original network. Second, Bayesian network (BN) models are established for each subnetwork, integrating multiple data sources and network structures. Finally, a joint Gibbs sampling mechanism, featuring forward and backward information flow, is designed to merge data from separate BN models and maximize the utilization of limited evidence, ensuring accurate outage location identification. In conclusion, the framework was validated on two modified public test systems, and comparative studies confirmed its effectiveness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Qualification of Digitized Legacy Fast Reactor Data

The Integral Fast Reactor (IFR) fuel compatibility test program (1984-1994) included a variety of fuel pin examinations conducted at the Hot Fuel Examination Facility (HFEF) and the Alpha-Gamma Hot Cell Facility (AGHCF). Hard copy data records of these examinations have been recovered, scanned, and preserved in PDF format. Many hard copy records are now qualified in accordance with an NRC-approved Quality Assurance Program Plan (QAPP), and there is an ongoing effort to qualify additional legacy records. This legacy fuel performance data is vital to support design and licensing of fast reactors with validation of state-of-the-art codes and advanced methods for design and analysis. Stakeholders can most easily utilize this data when the PDF scans have been converted into digital data tables. However, qualification of the scanned hard copy data does not qualify the digital data file resulting from the digitization of the data contained in the record; the subject matter expert (SME) must make a review of the digitized data table as well before it can be designated as qualified. This report outlines a peer review process to qualify the digital data file(s), typically in CSV format, corresponding to hard copy records in accordance with the existing QAPP.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗

Leveraging AI and Spatial Data to Unlock Pipeline Integrity Insights: NETL’s Advanced Infrastructure Integrity Model (AIIM)

Maintaining the integrity of natural gas infrastructure plays a critical role in ensuring energy security. Robust, data-driven foundational AI models for pipeline integrity can help address risk management and mitigation issues. Trusted foundational models can help with industry adoption and accelerate innovation by enhancing integrity predictions, reduce costs, and informing infrastructure build-out. The AIIM dashboard was released in 2022 and utilizes multi-ML models for ensemble-type insights. It was expanded to include analytics on reported incidents. It was developed as an ESRI Dashboard to support data visualization & interrogation and contains pipeline data and model results.

Advanced Infrastructure Integrity Model (AIIM)↗

Data and figures for "Integrated modeling of boron powder injection for real-time plasma-facing component conditioning"

This dataset contains raw and processed data, as well as supplementary figures used in the paper titled "Integrated modeling of boron powder injection for real-time plasma-facing component conditioning." The data includes simulation results for boron transport and deposition in DIII-D tokamak scenarios, and processed plots. It provides insights into the effects of boron powder injection on plasma-facing component conditioning and surface composition.

ablative particle injection↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Web-based Preprocessing and Visualization of 3D FIB Tomography Data for Nuclear Fuel Characterization

Three-dimensional (3D) focused ion beam (FIB) tomography enables reconstruction of internal nuclear fuel features that can't be fully evaluated through surface imaging alone. This capability supports characterization of fuel constituents and defects under thermal and irradiation conditions relevant to microreactor development. However, large tomography datasets can create data-handling, loading, and visualization challenges, especially when image-stack preparation and file conversion must be completed with separate tools. The Computational Ultraspatial Tomography Toolkit for High-Resolution Object Analysis Tools (CUTTRHOAT) is an open-source web application being developed to display FIB tomography datasets available through the Nuclear Research Data System (NRDS). The current alpha version requires prepared HDF5 datasets and has limited integrated data-preparation capabilities. This project improves CUTTHROAT by adding dataset-folder selection, automatic input detection, dataset scanning, missing-slice identification, blank-slice insertion, and image-stack-to-HDF5 conversion. Two applications will be compared: the baseline CUTTHROAT alpha workflow and the updated application containing the integrated data-handling and preprocessing functions. Evaluation will consider dataset detection accuracy, conversion success, loading time, rendering responsiveness, application stability, and user interaction. Preliminary results demonstrate successful loading of existing HDF5 files and converted image stacks, while testing also identified performance reductions caused by excessive blank-slice generation. The updated workflow reduces reliance on external preparation tools and supports more direct movement from image stacks to color-code 3D visualization. Future work includes refining missing-slice handling, integrating additional preprocessing functions, like a denoising feature, parsing TIFF metadata for automatic voxel scaling, and adding manual X, Y, and Z voxel-spacing inputs for PNG and JPEG.

36 - MATERIALS SCIENCE↗

Data for "Discovery, Characterization, and Application of Chromosomal Integration Sites in the Hyperthermophilic Archaeon Sulfolobus islandicus"

Sulfolobus islandicus , an emerging archaeal model organism, offers unique advantages for metabolic engineering and synthetic biology applications owing to its ability to thrive in extreme environments. Although several genetic tools have been established for this organism, the lack of well-characterized chromosomal integration sites has limited its potential as a cellular factory. Here, we systematically identified and characterized 13 artificial CRISPR RNAs targeting eight integration sites in S. islandicus using the CRISPR-COPIES pipeline and a multi-omics-informed computational workflow. We leveraged the endogenous CRISPR-Cas system to integrate the reporter gene lacS and validated heterologous expression through a β-galactosidase assay, revealing significant positional effects. As a proof of concept, we utilized these sites to genetically manipulate lipid ether composition by overexpressing glycerol dibiphytanyl glycerol tetraether (GDGT) ring synthase B (GrsB). This study expands the genetic toolbox for S. islandicus and advances its potential as a robust platform for archaeal synthetic biology and industrial biotechnology.

AI/ML↗

Deep Learning Advances Arctic River Water Temperature Predictions

The accelerated warming in the Arctic poses serious risks to freshwater ecosystems by altering streamflow and river thermal regimes. However, limited research on Arctic River water temperatures exists due to data scarcity and the absence of robust methodologies, which often focus on large, major river basins. To address this, we leveraged the newly released, extensive AKTEMP data set and advanced machine learning techniques to develop a Long Short-Term Memory (LSTM) model. By incorporating ERA5-Land reanalysis data and integrating physical understanding into data-driven processes, our model advanced river water temperature predictions in ungauged, snow- and permafrost-affected basins in Alaska. Our model outperformed existing approaches in high-latitude regions, achieving a median Nash-Sutcliffe Efficiency of 0.95 and root mean squared error of 1.0°C. The LSTM model learned air temperature, soil temperature, solar radiation, and thermal radiation—factors associated with energy balance—were the most important drivers of river temperature dynamics. Soil moisture and snow water equivalent were highlighted as critical factors representing key processes such as thawing, melting, and groundwater contributions. Glaciers and permafrost were also identified as important covariates, particularly in seasonal river water temperature predictions. Our LSTM model successfully captured the complex relationships between hydrometeorological factors and river water temperatures across varying timescales and hydrological conditions. This scalable and transferable approach can be potentially applied across the Arctic, offering valuable insights for future conservation and management efforts.

54 ENVIRONMENTAL SCIENCES↗

Radiative impact of record-breaking wildfires from integrated ground-based data

The radiative effects of wildfires have been traditionally estimated by models using radiative transfer calculations. Assessment of model-predicted radiative effects commonly involves information on observation-based aerosol optical properties. However, lack or incompleteness of this information for dense plumes generated by intense wildfires reduces substantially the applicability of this assessment. Here we introduce a novel method that provides additional observational constraints for such assessments using widely available ground-based measurements of shortwave and spectrally resolved irradiances and aerosol optical depth (AOD) in the visible and near-infrared spectral ranges. We apply our method to quantify the radiative impact of the record-breaking wildfires that occurred in the Western US in September 2020. For our quantification we use integrated ground-based data collected at the Atmospheric Measurements Laboratory in Richland, Washington, USA with a location frequently downwind of wildfires in the Western US. We demonstrate that remarkably dense plumes generated by these wildfires strongly reduced the solar surface irradiance (up to 70% or 450 Wm -2 for total shortwave flux) and almost completely masked the sun from view due to extremely large AOD (above 10 at 500 nm wavelength). We also demonstrate that the plume-induced radiative impact is comparable in magnitude with those produced by a violent volcano eruption occurred in the Western US in 1980 and continental cumuli.

54 ENVIRONMENTAL SCIENCES↗

Deliverable 6.7-Final Technical Report: Development Summary and Evaluation of the Solar Uncertainty Integrator (SUNI) Software

The Data Quality and Uncertainty Integration Project was a three-year effort to address stakeholder needs for assessing solar radiation resource data quality based on existing tools for estimating radiometer measurement uncertainties and assessing post-measurement data quality. The annual research objectives for the project addressed a logical progression of effort needed to achieve the ultimate project goal of developing the Solar Uncertainty Integrator (SUNI) software. This final technical report summarizes the development process for achieving these key research objectives and addresses the outreach and code development efforts in the final year of the project to develop a new solar irradiance data uncertainty integration software package.

14 SOLAR ENERGY↗

PV Degradation Modeling: Applying Geospatial Workflows with "PVDeg"

Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.

14 SOLAR ENERGY↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES↗

Data Center Cybersecurity, Supply Chain Risk Management, and Emerging Regulation Cohort Summary: Takeaways and Action Plans

This report summarizes the outcomes of the Data Center Cohort under the Department of Energy’s Technical Assistance for Digital Assurance (TADA) initiative, aimed at enhancing grid resilience through cybersecurity, supply chain risk management (SCRM), and Cyber-Informed Engineering (CIE). The cohort engaged 17 organizations across utilities, data center operators, vendors, and technology providers in three sessions combining presentations, discussions, and exercises. Key topics included AI-driven load behavior, cybersecurity vulnerabilities in UPS/BESS and cooling systems, governance gaps at utility–data center boundaries, and supply chain integrity. Five cross-cutting themes emerged: interconnection architecture vulnerabilities, fragmented governance, AI-driven stability risks, lack of regulatory frameworks, and long-term supply chain concerns. Actionable recommendations were developed, including implementing DMZ segmentation, formalizing vendor access agreements, designing AI workload limits, and advancing standards through NERC and state-level programs. These strategies aim to strengthen resilience, clarify responsibilities, and ensure secure integration of data centers into the grid.

24 - POWER TRANSMISSION AND DISTRIBUTION↗