Search NASA⌕ Search

SEARCH · Search NASA

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Towards Integrating Data Quality Assessments and Radiometer Uncertainty for Determining the Expanded Uncertainty of Three-Component Solar Radiation Measurements

Accurate solar irradiance data are fundamental for determining the design and performance characteristics of photovoltaic systems. The uncertainty of solar irradiance measurements depends on many factors including radiometer design, calibration, installation, maintenance, and operational environment. The key contributors to this uncertainty can be classified as the measurement uncertainty of a particular radiometer and the operational uncertainty determined for the time of measurement. Radiometer measurement uncertainty estimates (U R ) can be based on well-established methods used as part of the radiometer calibration process. Estimates of operational uncertainties (U o ) require consideration of additional site-specific factors that affect data quality. A method is needed for establishing the accuracy of solar irradiance data by integrating an existing data quality process and measurement uncertainty estimates for specific radiometers. An algorithm has been developed to integrate data quality analyses and measurement uncertainty estimates for three-component solar irradiance data: global horizontal (total hemispheric) irradiance, direct normal (beam) irradiance, and diffuse horizontal (sky) irradiance collected at one- to 60-minute intervals. The algorithm has been tested using one-minute irradiance measurements. The goal of the project is to distribute a user-friendly software package based on the new algorithm.

data integrity↗

An integrated data management and informatics framework for continuous drug product manufacturing processes: A case study on two pilot plants

The pharmaceutical industry continuously looks for ways to improve its development and manufacturing efficiency. In recent years, such efforts have been driven by the transition from batch to continuous manufacturing and digitalization in process development. To facilitate this transition, integrated data management and informatics tools need to be developed and implemented within the framework of Industry 4.0 technology. Here, in this regard, the work aims to guide the data integration development of continuous pharmaceutical manufacturing processes under the Industry 4.0 framework, improving digital maturity and enabling the development of digital twins. This paper demonstrates two instances where a data integration framework has been successfully employed in academic continuous pharmaceutical manufacturing pilot plants. Details of the integration structure and information flows are comprehensively showcased. Approaches to mitigate concerns in incorporating complex data streams, including integrating multiple process analytical technology tools and legacy equipment, connecting cloud data and simulation models, and safeguarding cyber-physical security, are discussed. Critical challenges and opportunities for practical considerations are highlighted.

59 BASIC BIOLOGICAL SCIENCES↗

Towards physics-inspired data-driven weather forecasting: integrating data assimilation with a deep spatial-transformer-based U-NET in a case study with ERA5

Abstract. There is growing interest in data-driven weather prediction (DDWP), e.g., using convolutional neural networks such as U-NET that are trained on data from models or reanalysis. Here, we propose three components, inspired by physics, to integrate with commonly used DDWP models in order to improve their forecast accuracy. These components are (1) a deep spatial transformer added to the latent space of U-NET to capture rotation and scaling transformation in the latent space for spatiotemporal data, (2) a data-assimilation (DA) algorithm to ingest noisy observations and improve the initial conditions for next forecasts, and (3) a multi-time-step algorithm, which combines forecasts from DDWP models with different time steps through DA, improving the accuracy of forecasts at short intervals. To show the benefit and feasibility of each component, we use geopotential height at 500 hPa (Z500) from ERA5 reanalysis and examine the short-term forecast accuracy of specific setups of the DDWP framework. Results show that the spatial-transformer-based U-NET (U-STN) clearly outperforms the U-NET, e.g., improving the forecast skill by 45 %. Using a sigma-point ensemble Kalman (SPEnKF) algorithm for DA and U-STN as the forward model, we show that stable, accurate DA cycles are achieved even with high observation noise. This DDWP+DA framework substantially benefits from large (O(1000)) ensembles that are inexpensively generated with the data-driven forward model in each DA cycle. The multi-time-step DDWP+DA framework also shows promise; for example, it reduces the average error by factors of 2–3. These results show the benefits and feasibility of these three components, which are flexible and can be used in a variety of DDWP setups. Furthermore, while here we focus on weather forecasting, the three components can be readily adopted for other parts of the Earth system, such as ocean and land, for which there is a rapid growth of data and need for forecast and assimilation.

54 ENVIRONMENTAL SCIENCES↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Enabling Data Exchange and Data Integration with the Common Information Model: An Introduction for Power Systems Engineers and Application Developers

The Common Information Model (CIM) is an open-source information model that is used to model an electrical network and the various equipment used on the network. CIM is widely used for data exchange of bulk transmission power systems and is finding increasing use for distribution systems. Use of a non-proprietary information model (such as CIM) that has been agreed upon and adopted by numerous utilities, vendors, and researchers allows significant reduction in the effort and cost of data integration. Likewise, adoption of open data platforms built around the CIM increases available functionalities for managing and optimizing the smart grid of the future. This report is intended as an introduction to CIM for utility engineers, power systems researchers, and application developers, providing a broad view of the CIM and how particular profiles can be adapted for various use cases. Unlike most other CIM introduction documents and the International Electrotechnical Commission (IEC) standards (which are mostly targeted to an audience of data scientists, enterprise database managers, and platform developers), this report is intended for users of traditional power systems analysis software and other readers without any prior experience with canonical information models, data profiles, or UML modeling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Impact Analysis of Data Integrity Attacks on FACTS-based Wide-Area Voltage Control System

Energy management system (EMS) consists of several wide-area control applications that serve as a backbone for security, stability, and reliability of the power system. Wide-area voltage control system (WAVCS), one of the critical wide-area applications, operates in coordination with local Flexible AC Transmission System (FACTS) devices to provide voltage security and optimal management of active and reactive power resources. Since the WAVCS relies on wide-area communication and data sharing devices, possible cybersecurity vulnerabilities have to be addressed to ensure the closed-loop operation of WAVCS. In this paper, we present a methodology for performing an impact analysis of cyber-attacks in WAVCS cybersecurity. In particular, different types of data integrity attacks, such as malicious tripping, fault replay, and signal altering attacks, are considered, and detailed impact analysis is conducted in a testbed environment using the Kundur's four machine two-area system. For performing an impact analysis, the transient voltage stability of the sensitive bus voltage is studied, followed by the quantitative assessment and severity ranking using the voltage profile index. Our experimental evaluation reveals that the data integrity attacks on control signals exhibit a higher attack severity than on the measurement signals. Further, the severity of these attacks varies with nature (static or dynamic), location, and types of attacks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Low Yield Nuclear Monitoring Physics Experiment 1 – Integrated Data Acquisition System Design and Initial Observations

The report documents the design of the Integrated Data AcQuisition (IDAQ) system and observations recorded during the first in a series of underground chemical explosions conducted on the Nevada National Security Site (NNSS) in southern Nevada. Experiments are funded as part of Low Yield Nuclear Monitoring (LYNM) research and development within the United States National Nuclear Security Administration NA-22 nuclear non-proliferation program. The series is part of the broader Physical Experiment 1 (PE1) being conducted in and around the P-tunnel facility on the NNSS. Each explosive experiment utilizes several tons of comp-B to generate signals recorded by a broad suite of instrumentation. The IDAQ serves as the backbone for all subsurface instrumentation providing precise time synchronization, remote control, data exfiltration and backup, along with recording several sensing modalities throughout the underground complex that includes ground motion, environmental conditions, and electromagnetic signals.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Development of a Framework for Data Integration, Assimilation, and Learning for Geological Carbon Sequestration (DIAL-GCS) (Final Report)

This project aimed to develop and demonstrate a Data Integration, Assimilation, and Learning framework for geologic carbon sequestration projects (DIAL-GCS). DIAL-GCS is an intelligence monitoring system (IMS) for automating GCS closed-loop management by leveraging recent developments in machine learning technologies, complex event processing (CEP), and reduced-order modeling. The safe and efficient operation of GCS repositories requires integrated monitoring to track the injected CO¬2 as it moves within a storage reservoir. GCS projects are data intensive, as a result of proliferation of digital instrumentation and smart-sensing technologies. GCS projects are also resource intensive, often requiring multidisciplinary teams performing different monitoring, verification, accounting (MVA) tasks throughout the lifecycle of a project to ensure secure containment of injected CO2. The success of GCS thus depends in a large part on our ability to access, assimilate, and analyze heterogeneous data and information sources in a timely manner. This project included a number of meaningful and necessary tasks to transform the human domain knowledge into machine-interpretable rules for automating knowledge extraction and discovery in GCS. The specific technical objectives of the proposed DIAL-GCS project were to develop an ontology-driven GCS data management module for storing, querying, and exchanging GCS data (both historic and live sensor data) from multiple sources and in heterogeneous formats. Incorporate a CEP engine for detecting abnormal situations by seamlessly combining expert knowledge, rule-based reasoning, and machine learning. Enable uncertainty quantification and predictive analytics using a combination of coupled-process modeling, AI/ML methods, and reduced-order modeling, and integrate and demonstrate the system’s capabilities with both real and simulated data. As far as we know, this is one of the first projects aimed to develop intelligent monitoring systems (IMS) targeting the GCS. Under this project, the team had developed a large number of web applications and scientific algorithms that contribute the main theme of intelligent monitoring. The team has published more than a dozen peer reviewed papers and disseminated the research results at multiple technical meetings.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗