Search NASASearch

SEARCH · Search NASA

Results for “Data Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Global Archaeal Diversity Revealed Through Massive Data Integration: Uncovering Just Tip of Iceberg

The domain of Archaea has gathered significant interest for its ecological and biotechnological potential and its role in helping us to understand the evolutionary history of Eukaryotes. In comparison to the bacterial domain, the number of adequately described members in Archaea is relatively low, with less than 1000 species described. It is not clear whether this is solely due to the cultivation difficulty of its members or, indeed, the domain is characterized by evolutionary constraints that keep the number of species relatively low. Based on molecular evidence that bypasses the difficulties of formal cultivation and characterization, several novel clades have been proposed, enabling insights into their metabolism and physiology. Given the extent of global sampling and sequencing efforts, it is now possible and meaningful to question the magnitude of global archaeal diversity based on molecular evidence. To do so, we extracted all sequences classified as Archaea from 500 thousand amplicon samples available in public repositories. After processing through our highly conservative pipeline, we named this comprehensive resource the ‘Global Archaea Diversity’ (GAD), which encompassed nearly 3 million molecular species clusters at 97% similarity, and organized it into over 500 thousand genera and nearly 100 thousand families. Saline environments have contributed the most to the novel taxa of this previously unseen diversity. The majority of those 16S rRNA gene sequence fragments were verified by matches in metagenomic datasets from IMG/M. These findings reveal a vast and previously overlooked diversity within the Archaea, offering insights into their ecological roles and evolutionary importance while establishing a foundation for the future study and characterization of this intriguing domain of life.

59 BASIC BIOLOGICAL SCIENCES

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES

Integral Nuclear Data and Benchmarking Needs for Fusion Energy Systems

Fusion energy systems are currently being designed and optimized using radiation transport codes. To deal with the unique environment inside a fusion-based system, many of these designs incorporate novel materials able to withstand the high radiation fields, ensure adequate cooling and thermal protection, and produce tritium. Validation plays a vital role in building trust in the predictive power of these models and computational methods. Validation of a code consists of modeling documented real-world experiments and comparing the code-predicted response to the measured response. Adequate validation requires measured responses from real-world experiments, also known as integral data, that mimic the system being designed, including materials, impinging radiation, and temperature, among other variables. The most trusted integral data are experimental responses that have been through a rigorous benchmarking process that develops a recommended computational model and evaluates all experimental uncertainties. Finally, there are a few research groups around the world that have been producing integral data for fusion applications, but a substantial investment is needed to address the unique validation needs of the fusion community.

Fusion

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES

Ensemble Federated Machine Learning‐Based Cybersecurity Situational Awareness in Microgrid Network

Cyber-physical microgrids are vulnerable to stealthy cybersecurity threats that disguise their actions through the exploitation of system knowledge. Such actions can severely impacts microgrids deployed in defense bases, slowing the response time of military forces during national emergencies. Several machine-learning algorithms have been proposed to detect intrusions in the grid networks; however, these traditional machine-learning algorithms lack data privacy and are subject to several adversarial machine-learning threats. This paper proposes a novel federated machine learning (FML)-based three-model framework to detect and identify stealthy data-integrity attacks while ensuring data privacy in microgrid networks. The proposed architecture uses a variational mode decomposition technique to extract derived features from incoming measurement and control datasets. The extraction of these derived features allows FML models to learn minute variations in data patterns that allow them to perform significantly better than the models trained with generic datasets consisting of raw features. Our experimental results show the efficient performance of the proposed methodology against different types of data integrity attacks while considering primary and secondary controllers in microgrids. Further, the applied FML-integrated random forest ensemble algorithm outperforms the existing generic FML algorithms during noisy and noise-free datasets with prediction latencies of only 91–134 µs per sample within the 0.1 s sampling interval and requires communication bandwidth of around ∼8.25 KB/s at the control center and ∼2.7 KB/s per edge client for communication.

24 POWER TRANSMISSION AND DISTRIBUTION

Comparability of Liquid Chromatography Tandem Mass Spectrometry Analysis of Dissolved Organic Matter across Laboratories

Non-targeted liquid chromatography tandem highresolution mass spectrometry (LC−MS/MS) is increasingly applied for the structure-resolved chemical analysis of dissolved organic matter (DOM). With new developments in MS instrumentation and analysis software, the approach has gained substantial momentum over the past decade. However, achieving high-quality analytical data that is reproducible and comparable across laboratories can be a bottleneck in non-targeted metabolomics and organic matter chemical analysis, especially for data reuse in repository-scale analyses. Understanding the capabilities as well as challenges of comparing LC−MS/MS data from different laboratories is necessary for inferring global trends from public data sets. To illuminate instrumentation factors that drive differences and variability, we used a standardized data analysis pipeline, including classical (CMN) and featurebased molecular networking (FBMN), to analyze data from a ring trial by 24 laboratories on identical sample sets of algal and DOM extracts that were mixed in predefined concentrations and spiked with standards. Our results showed that data sets from similar mass spectrometer types with unified instrument parameters were qualitatively comparable, resolving the same general trends and shared mass spectral features. Interlaboratory comparability was best for high-intensity features, while low-intensity features showed greater detection variability. Our analysis also highlights challenges when comparing data from instruments with different acquisition rates or operating with less standardized methods. Lastly, we provide recommendations for data integration, public data sharing, standardization, and best practices for standardized LC−MS/MS data acquisition, which will be critical for long-term time series and intercomparability of DOM chemical analyses.

DOM

Cloud-based Testbed for Adaptive Under-Frequency Load Shedding with High DER Penetration

Increasing penetration of distributed energy resources and behind-the-meter renewables may soon disrupt the efficacy of critical protection schemes, such as under-frequency load shedding (UFLS). Improved data exchange and coordination across the transmission-distribution boundary will be required to maintain reliability of bulk electric system. Standards-based data integration platforms using agreed-upon semantic vocabularies, such as the Common Information Model, will be key to enabling adaptive protection schemes requiring synthesized data from both the bulk power system and behind-the-meter resources. This paper introduces a cloud-based open-source data integration environment and UFLS clustering algorithm being developed to enable adaptive relay coordination between transmission and distribution utilities in the state of Vermont.

Anderson, Alexander A.

Regression Analysis with the Directed Infusion of Data

Integrating artificial intelligence and machine learning tools into industry necessitates large-scale collaborative efforts that ensure the robust and accurate execution of downstream analytics such as time series prediction, uncertainty quantification, grid optimization, and condition monitoring. However, concerns related to data privacy pervade the nuclear industry due to the proprietary nature of its data and the possibility of data leakage. Legacy techniques such as encryption often require the explicit transmission of data to trustworthy parties, thereby inviting data leakage concerns. The ideal collaboration scenario avoids the explicit dissemination of data/code while maintaining experimental fidelity, which is currently accomplished using various techniques such as trusted execution environments, homomorphic encryption, differential privacy, and multimatrix masking. These techniques, however, often necessitate a trade-off between trust, efficiency, and utility. This article extends a previously proposed technique called the directed infusion of data (DIOD) that ensures data privacy, allows for scalable obfuscation, and combats the risk of data leakage without compromising utility. The experiments discussed in this article examine a regression-type scenario using DIOD with the goal of preserving the inferential link between two variables. Using the point-kinetics equations, regression experiments compare the performance of a model trained using the original data to that of a model trained using the obfuscated data, which produced identical results. Our claim is further strengthened by an information theoretic proof and experiment, which showed that the inferential content between variables remains the same after obfuscation, thereby avoiding the required communication of the proprietary data.

47 - OTHER INSTRUMENTATION

Captan+X Data Converter Integration

Fermi National Accelerator Laboratory's CAPTAN (Compact And Programmable daTa Acquisition Node) series provides a flexible hardware platform for data acquisition across a range of experiments and facilities. The latest iteration, CAPTAN+X, is built around a Kintex-7 FPGA supporting four FPGA Mezzanine Card (FMC) connections. As part of a broader laboratory effort to bring facility systems under a Model-Based Systems Engineering (MBSE) framework, CAPTAN+X is one of several systems slated to be incorporated into this modeling environment in the near term. A necessary step toward that goal is incorporating the platform's core functionality, which centers on integration with the LXD31K4 FMC, a data converter module combining dual AD9652 analog-to-digital converters and dual AD9142A digital-to-analog converters. Achieving compatibility required resolving pin-mapping conflicts between the LXD31K4's High Pin Count connector and the CAPTAN+X's available pin types, adapting a Board Support Project originally written for an UltraScale-class evaluation board to the Kintex-7 architecture, replacing incompatible primitives, restructuring clock distribution, and manually configuring chip initialization in place of an unsupported soft-processor-based approach. Functional verification of the ADC and DAC channels, followed by closed-loop testing combining both converters with real-time filtering, confirmed correct operation of the integrated system. These results establish a working hardware and firmware baseline for the CAPTAN+X platform, positioning it for future inclusion in the laboratory's growing MBSE modeling effort.

Espinoza, David [Illinois U., Urbana (main)]

CAPTAN+X Data Converter Integration

Fermi National Accelerator Laboratory's CAPTAN (Compact And Programmable daTa Acquisition Node) series provides a flexible hardware platform for data acquisition across a range of experiments and facilities. The latest iteration, CAPTAN+X, is built around a Kintex-7 FPGA supporting four FPGA Mezzanine Card (FMC) connections. As part of a broader laboratory effort to bring facility systems under a Model-Based Systems Engineering (MBSE) framework, CAPTAN+X is one of several systems slated to be incorporated into this modeling environment in the near term. A necessary step toward that goal is incorporating the platform's core functionality, which centers on integration with the LXD31K4 FMC, a data converter module combining dual AD9652 analog-to-digital converters and dual AD9142A digital-to-analog converters. Achieving compatibility required resolving pin-mapping conflicts between the LXD31K4's High Pin Count connector and the CAPTAN+X's available pin types, adapting a Board Support Project originally written for an UltraScale-class evaluation board to the Kintex-7 architecture, replacing incompatible primitives, restructuring clock distribution, and manually configuring chip initialization in place of an unsupported soft-processor-based approach. Functional verification of the ADC and DAC channels, followed by closed-loop testing combining both converters with real-time filtering, confirmed correct operation of the integrated system. These results establish a working hardware and firmware baseline for the CAPTAN+X platform, positioning it for future inclusion in the laboratory's growing MBSE modeling effort.

Espinoza, David [Illinois U., Urbana (main)]

Radiation-induced bowing of SiC/SiC composites under neutron flux gradients—integral experimental data for model validation

Here, the radiation-induced swelling of SiC and its composites, including strong dependencies on temperature and dose, can drive significant lateral bowing in the presence of temperature and/or dose gradients. In recent years, simulations have been performed to assess the extent of bowing in SiC composite light-water reactor (LWR) fuel cladding and boiling water reactor (BWR) channel boxes. However, to date, no integral experimental data exist to validate these models. This work provides the first experimental bowing evaluation of three ∼380 mm long SiC composite specimens irradiated under varying neutron dose gradients (∼50°C–60°C, 0.03–0.06 dpa): two tubes (∼9.8 mm diameter) and a miniature BWR channel box (∼30 mm square). The measured radiation-induced length swelling (∼0.3%–0.7% linear) was consistently 10%–21% higher than values obtained from 3D finite element structural analyses with inputs from 3D radiation transport calculations. This discrepancy could be at least partially explained by differences in dose rate (∼10 -8 dpa/s) compared to the literature data (∼10-6 dpa/s) used to establish the dose-to-swelling correlations in the model. Nevertheless, the modeled bowing magnitudes (<2 mm) obtained from finite element analyses and simple analytical equations were within the bounds of the experimental measurements for all specimens. With improved confidence in the ability to predict the structural response and measure the macroscopic deformations, future experiments will target transient bowing under neutron flux gradients at representative LWR temperatures and assess whether grid spacers can mitigate the tens of millimeters of bowing that would otherwise be expected in ∼4 m long LWR components.

bowing

Ranking Biological Features in Soil-Based Microbial Multi-Omics Data with Integration Modeling

Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.

54 ENVIRONMENTAL SCIENCES

Integrating AI Data Centers with the Power Grid

The rapid expansion of artificial intelligence (AI) has triggered an unprecedented surge in electricity demand, with US data center energy use projected to double or triple 2023 levels by 2028. This exponential growth places strain on grid infrastructure, which can hinder timely construction of desired computing capacity. To bridge this supply-demand gap, utilities and AI developers are increasingly turning to demand flexibility, a strategy that incentivizes shifting or reducing power use during peak periods of grid stress. Data centers are uniquely equipped for flexible operations due to their digital workloads, built-in redundancy, and onsite energy assets. This article outlines four primary mechanisms to enable data center flexibility: computational load flexibility (shifting tasks temporally or geographically), flexible use of core facility infrastructure adjustments, energy storage utilization, and onsite electricity generation. To encourage adoption, utilities are deploying new tariff designs, including voluntary interruptible service riders, mandated flexibility requirements, and streamlined interconnection processes for flexible loads. For the highly capitalized and rapidly growing AI industry, the primary motivators for embracing these strategies are expediting facility interconnection, satisfying emerging regulatory mandates, and mitigating community resistance. While demand flexibility cannot substitute the long-term need for new bulk power generation, it serves as an essential, immediate solution for enabling near-term deployment. By transforming data centers from grid stressors into stabilizing assets, flexible operations can ensure reliable grid integration, ease market pressures, and support a resilient power system.

24 POWER TRANSMISSION AND DISTRIBUTION