Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Nuclear data sensitivity and uncertainty study of copper-reflected integral experiments [Slides]

This presentation touches on reducing uncertainties in intermediate-energy actinide nuclear data and this continues to be a high priority for many applications. The goal of PARADIGM (PARallel Approach of Differential and InteGral Measurements): accelerate efforts to reduce biases and uncertainties in nuclear data through improvements to the nuclear data pipeline. This presentation includes integral experiments and a summarization of existing copper nuclear data.

Cu63↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Building partnerships for development of sustainable energy systems with atmospheric measurements

Atmospheric dynamics often play a critical role in the sustainability and reliability of diverse forms of energy production. This is especially true for the growing number of renewable energy deployments that harness aspects of the environment for power production. While the University of Memphis has a strong research background in energy systems, we have little experience working with the Earth and Environmental Systems Science Division (EESSD) and their associated User Facilities. Of particular interest to us is the Atmospheric Science Research and the Atmospheric Radiation Measurement (ARM) user facility to address surface-boundary layer interactions and physical phenomena. One of the major challenges for understanding and developing energy systems and management platforms is accurate modeling/forecasting of atmospheric conditions across disparate spatial and temporal scales. These conditions are often required to understand the lowest levels of the atmospheric boundary layer, but are also important to understand higher atmospheric conditions where aerosols affect cloud development. The objective of this work was to develop partnerships with national laboratories for collaboration on environmental science and its intersection with sustainable energy systems, as well as to leverage the ARM user facility data repositories to enhance our research capabilities in energy systems and their inter-dependence on environmental systems for future engagement with EESSD. Specifically, we accomplished these objectives by (1) developing collaborations with Oakridge National Laboratory ARM Data Science and Integration Group which resulted in student internships, (2) employed ARM data to develope modeling of the atmospheric boundary layer optical turbulence, and (3) optimally-sized large-scale renewable energy systems and their associated energy storage systems with ARM repository data.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Development of a Distribution Optimal Power Flow Federate for Open-Source OEDI-SI Platform

Increasing numbers of distributed generators in the electric power distribution networks require developing a control strategy to optimize solutions in real time. Linearized optimal distribution flow development has seen growth and acceptance in the distribution systems literature for efficiently modeling the \glspl{opf} for distribution systems. This paper examines the implementation and integration procedure for linearized optimal distribution flow federate to \gls{oedisi} platform. Specifically, we discuss i) the usage of the \gls{oedisi} platform, ii) obtaining a tractable solution using developed \gls{opf} federate, and iii) validation of solutions and bench-marking the \gls{oedisi} platform with developed \gls{opf} federate using OpenDSS. In brief, we demonstrate how a general linearized optimal distribution flow federate can be developed and integrated with a co-simulation environment to mimic real-world examples. The efficacy of the proposed method is demonstrated using the IEEE 123-bus test system under different scenarios to obtain a tractable solution and compare its results.

Sadnan, Rabayet↗

Analyzing the Impact of Future Weather Data on Energy Consumption in Weatherization Assistant

This study supports the mission of the U.S. Department of Energy’s Weatherization Assistance Program (WAP), which aims to increase the energy efficiency of dwellings and reduce their total residential expenditures. Specifically, we examine how projected future climate conditions may affect residential building energy performance by integrating future weather data into the National Energy Audit Tool (NEAT). Since WAP evaluates the cost-effectiveness of retrofit measures over lifespans of up to 30 years, accounting for evolving climate conditions is increasingly important. To reflect future household energy demands, this study replaces historically based Typical Meteorological Year (TMY3) weather inputs with Future Typical Meteorological Year (fTMY) datasets derived from global climate model (GCM) projections. A simulation-based framework was established to enable NEAT analysis under future weather conditions. This workflow involves converting EPW-format weather files into JSON inputs compatible with NEAT and generating degree-hour metrics needed for load calculations. The fTMY dataset used in this study was developed by Oak Ridge National Laboratory through downscaling of six GCMs under different emission scenarios and covers the period from 2020 to 2100. In contrast, the TMY3 dataset is based on historical weather data from 1961 to 1990. Simulations were conducted for benchmark single-family prototype buildings across ASHRAE climate zones 1–7, which cover all regions of the U.S. except the subarctic Zone 8 in northern Alaska, evaluating both heating and cooling loads under TMY3 and fTMY conditions. Four foundation types were tested, while heating systems were standardized, as NEAT does not differentiate thermal energy load by HVAC system type in its load calculations. Results show that fTMY weather input consistently yield lower heating loads and higher cooling loads across most locations, aligning with expected climate warming trends. Notably, colder regions such as zones 6A, 6B, and 7 experience marked reductions in heating load, while warmer and transitional zones, such as 2A (Lufkin, TX) and 3C (San Francisco, CA), have substantial increases in cooling loads. Although this study does not directly assess the performance of retrofit measures under future climate conditions, it provides a critical foundation for doing so. By quantifying shifts in baseline (i.e., pre-retrofit case) energy loads between historical and future weather files, the study highlights the importance of integrating climate-responsive data into audit tools. These findings will inform future efforts to evaluate the long-term effectiveness and cost-effectiveness of weatherization measures under changing climate conditions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A U.S. Scientific Community Vision for Sustained Earth Observations of Greenhouse Gases to Support Local to Global Action

Managing carbon stocks in the land, ocean, and atmosphere under changing climate requires a globally‐integrated view of carbon cycle processes at local and regional scales. The growing Earth Observation (EO) record is the backbone of this multi‐scale system, providing local information with discrete coverage from surface measurements and regional information at global scale from satellites. Carbon flux information, anchored by inverse estimates from spaceborne Greenhouse Gas (GHG) concentrations, provides an important top‐down view of carbon emissions and sinks, but currently lacks global continuity at assessment and management scales (<100 km). Partial‐column data can help separate signals in the boundary layer from the overlying atmosphere, providing an opportunity to enhance surface sensitivity and bring flux resolution down from that of column‐integrated data (100–500 km). Based on a workshop held in September 2024, the carbon cycle community envisions a carbon observation system leveraging GHG partial columns in the lower and upper troposphere to weave together information across scales from surface and satellite EO data, and integration of top‐down/bottom‐up analyses to link process understanding to global assessment.

Parazoo, Nicholas C. [California Institute of Tech↗

DE-FE0029488 - North Dakota Integrated Carbon Capture and Storage Complex Feasibility Study Public Data

Data from award DE-FE0029488 - North Dakota Integrated Carbon Capture and Storage Complex Feasibility Study performed by the Energy & Environmental Research Center including the following: - 2D Seismic {Input data, sgy files, maps, logs, and descriptors} - Core Petrophysics {Core analysis of plugs from the two stratigraphic test wells (Flemmer-1 [API 33-057-00039] and BNI-1 [API 33-065-00018])} - North Dakota Oil and Gas File No 37380 Files - North Dakota Oil and Gas File No 37672 Files - Well Testing Data {Summary of well testing methods and results from the stratigraphic test wells (Flemmer-1 and BNI-1)} Additional References: https://www.netl.doe.gov/sites/default/files/2017-12/Wesley-Peck-_Mastering-the-Subsurface_CarbonSAFE-Phase-II_August-2017-final.pdf Peck, W.D., Ayash, S.C., Klapperich, R.J., Gorecki, C.D. (2019) The North Dakota integrated carbon storage complex feasibility study, International Journal of Greenhouse Gas Control, Volume 84, 2019, Pages 47-53, https://doi.org/10.1016/j.ijggc.2019.03.001

Carbon Storage↗

Streaming Readout and Data-Stream Processing With ERSAP

With the exponential growth in the volume and complexity of data generated at high-energy physics and nuclear physics research facilities, there is an imperative demand for innovative strategies to process this data in real or near-real-time. Given the surge in the requirement for high-performance computing, it becomes pivotal to reassess the adaptability of current data processing architectures in integrating new technologies and managing streaming data. This paper introduces the ERSAP framework, a modern solution that synergizes flow-based programming with the reactive actor model, paving the way for distributed, reactive, and high performance in data stream processing applications. Additionally, we unveil a novel algorithm focused on time-based clustering and event identification in data streams. The efficacy of this approach is further exemplified through the data-stream processing outcomes obtained from the recent beam tests of the EIC prototype calorimeter at DESY.

Vardan, Gyurjyan↗

Enhancing fire emissions inventories for acute health effects studies: integrating high spatial and temporal resolution data

Daily fire progression information is crucial for public health studies that examine the relationship between population-level smoke exposures and subsequent health events. Issues with remote sensing used in fire emissions inventories (FEI) lead to the possibility of missed exposures that impact the results of acute health effects studies. This paper provides a method for improving an FEI dataset with readily available information to create a more robust dataset with daily fire progression. High temporal and spatial resolution burned area information from two FEI products are combined into a single dataset, and a linear regression model fills gaps in daily fire progression. The combined dataset provides up to 71% more PM 2.5 emissions, 69% more burned area, and 367% more fire days per year than using a single source of burned area information. The FEI combination method results in improved FEI information with no gaps in daily fire emissions estimates. The combined dataset provides a functional improvement to FEI data that can be achieved with currently available data.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Accuracy of kinetic equilibrium reconstruction of NSTX and NSTX-U plasmas and its impact on the transport and stability analysis

An accurate magnetohydrodynamic (MHD) equilibrium reconstruction is an essential starting point for stability and transport plasma analysis. Herein this work describes an approach for obtaining kinetic equilibrium reconstructions using the OMFIT framework, which has been applied for the first time to spherical tokamak data from NSTX and NSTX-U. The EFIT equilibrium solver is integrated with experimental data analysis procedures and subsequent TRANSP transport simulations to enhance the accuracy of the reconstruction, in particular, at the edge region, by adding constraints on the total pressure and current density profiles, based on the transport code solution. The accuracy of the equilibrium reconstruction depends on the uncertainty and number of constraints, as well as the choice of basis functions to represent the pressure and current density profiles. Improved fidelity of the equilibrium reconstruction is demonstrated by reducing the variability of the magnetic axis and boundary locations from several centimeters, for reconstructions based on magnetic and experimental pressure constraints, to only several millimeters, for kinetic reconstructions based on transport code constraints, when different representations of basis functions were tested. The variability of the safety factor on axis was reduced ten times in the same sensitivity study. The accuracy of the equilibrium reconstruction and subsequent mapping of the experimental kinetic profile data have a significant impact on the trapped gyro Landau fluid and linear CGYRO turbulence simulations, which predict different spectra of unstable modes and turbulent fluxes for cases with different numbers of constraints in the equilibrium reconstruction. Conversely, the stability analysis performed using the GATO code shows plasmas that are stable to n = 1 MHD modes in both equilibria using magnetic and experimental pressure constraints as well as the transport code constrained equilibrium. However, a scan of parameters away from these conditions shows considerable deviation in the threshold of unstable modes between these reconstructions. Therefore, for reliable plasma analysis and use in turbulence and stability calculations, a high-fidelity equilibrium reconstruction with accurate kinetic constraints based on transport code solutions is necessary.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

2015-2017 California Vehicle Survey

The 2015-2017 California Vehicle Survey of residential and commercial light-duty vehicle owners in California assessed consumer preferences for vehicles and included a targeted sample of plug-in electric vehicle (PEV) owners. Resource Systems Group conducted the survey on behalf of the California Energy Commission. In addition to economic and demographic data, the survey integrated light-duty vehicle holding and use information with vehicle choice data collected via the stated preferences survey's set of eight vehicle and fuel type choice exercises. The PEV owner survey participants provided additional data on charging behavior, electricity rates, and their main motivations for purchasing PEVs.

1Hz data↗

2019 California Vehicle Survey

The 2019 California Vehicle Survey of residential and commercial light-duty fleet owners in California assessed consumer preferences for vehicles and included a targeted sample of plug-in electric vehicle (PEV) owners. In addition to economic and demographic data, the survey integrated light-duty vehicle holding and use information with vehicle choice data, which was collected via a set of eight exercises on vehicle and fuel type choice. The PEV owner survey participants provided additional data on charging behavior, electricity rates, and their main motivations for purchasing PEVs.

1Hz data↗

Integration of Condition-Based, Diagnostic, Prognostic, And Anomaly Detection Data into Reliability Models to Support a Predictive Maintenance Context

Reliability data employed in plant reliability models are an approximated integral representation of the past industrywide operational experience, and they neglect the present asset health status (available, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Ideally, in a predictive maintenance context, system reliability models should support decision making by propagating actual health information from the asset to the system level in order to provide a quantitative snapshot of system health and identify the most critical assets. Asset health should be informed solely by that specific asset’s current and historical performance data and should not be an approximated integral representation of the past industrywide operational experience (as currently performed by system reliability models through Bayesian updating processes). This paper proposes a reliability modeling approach that relies on asset diagnostic and prognostic assessments, along with monitoring data to measure asset health. We show how state-of-the art condition-based, diagnostic, prognostic, and anomaly detection models can be linked to system reliability models not in probability terms, but in terms of margin where margin is defined as the “distance” between the present status and an undesired event (e.g., failure or unacceptable performance). Then, we show how the propagation of margin data from the asset to the system level is performed through classical reliability models such as fault trees or reliability block diagrams. The described method is in fact able to propagate heterogenous health data from the asset to the system level in order to analytically assess system health.

97 MATHEMATICS AND COMPUTING↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗