Search NASASearch

SEARCH · Search NASA

Results for “Data Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Data Centers Gap Analysis [Slides]

Data centers and other large loads are a significant driver of unprecedented, near-term demand growth in the United States. Power system planners, utilities, regulators, and other stakeholders are grappling with how to integrate data centers on the system without comprising reliability, resiliency, and energy affordability. NLR is pursuing work to develop a siting and decision-making tool that would draw on power systems modeling expertise to achieve granular representation of trade-offs involved in data center sitting and development. This slide deck supports the same workstream by reviewing the literature to identify mitigation options to facilitate near-term integration of large loads and by presenting options for pursuing data development and/or modeling projects to improve representation of siting options.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Designing a User Interface for Real-Time Magnetometer Data Acquisition

The Matter-wave Atomic Gradiometer Interferometric Sensor (MAGIS-100) is a next-generation quantum sensor designed to search for ultralight dark matter and explore new frontiers in quantum mechanics. Due to the experiment s sensitivity to magnetic interference, a magnetometer trolley system was developed to scan magnetic fields along a vacuum tube. Interacting with the system required command-line inputs, creating usability challenges. To improve accessibility and streamline data acquisition, I developed a graphical user interface (GUI) using Python and the customtkinter library. The GUI supports real-time data display, state/mode switching, command execution, and CSV file management. I collaborated with another intern to integrate data visualization features into the GUI, allowing users to generate 3D plots of post-acquisition magnetic field data. In the future, I aim to fix the real-time plotting feature as it results in an unresponsive GUI.

Mendez, Milagros [DuPage Coll.]

Designing a User Interface for Real-Time Magnetometer Data Acquisition

The Matter-wave Atomic Gradiometer Interferometric Sensor (MAGIS-100) is a next-generation quantum sensor designed to search for ultralight dark matter and explore new frontiers in quantum mechanics. Due to the experiment’s sensitivity to magnetic interference, a magnetometer trolley system was developed to scan magnetic fields along a vacuum tube. Interacting with the system required command-line inputs, creating usability challenges. To improve accessibility and streamline data acquisition, I developed a graphical user interface (GUI) using Python and the customtkinter library. The GUI supports real-time data display, state/mode switching, command execution, and CSV file management. I collaborated with another intern to integrate data visualization features into the GUI, allowing users to generate 3D plots of post-acquisition magnetic field data. In the future, I aim to fix the real-time plotting feature as it results in an unresponsive GUI.

Mendez, Milagros [DuPage Coll.]

Artificial Intelligence for Enhancing Multiscale Analysis: Buildings Focus

This project aims to develop multi-scale building energy data, potentially improving the representation of the U.S. buildings sector in GCAM-USA, an U.S.-focused human-energy-Earth systems model. Existing building energy datasets are typically limited to national or regional levels, which constrains the ability of models to capture fine-scale human-energy-Earth systems interactions and reduces their relevance for decision-making on issues such as energy security, resilience, and energy planning. By leveraging AI and advanced data integration methods, this work fuses multiple existing datasets to enhance the physical and geographic representation of both residential and commercial building energy use. So far, progress includes processing residential building data, designing the data structure for commercial buildings, and testing AI approaches for integrating datasets and addressing spatial-temporal gaps. This effort can not only advances GCAM-USA’s capability in modeling the buildings sector but also supports broader DOE missions, such as developing digital testbeds, enhancing grid resilience analysis, and improving building–energy system modeling at decision-relevant scales.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Workshop: Advanced Metering for Decarbonization: Electric Vehicles and 24/7 Carbon-Free Electricity

The purpose of this workshop is to learn more about advanced metering best practices for meeting the goals of EO 14057. This will be a 2-part session, first part to include presentations on the FEMP best practice work related to metering: 1) Electric Vehicles (EVs) and EV charging station electricity use. 2) Integrating data sources to calculate hourly carbon pollution-free electricity (CFE). Second part will facilitate small group discussions with a problem-solving activity.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC

Extinction Monitoring of Pulsed Proton Beams Using FPGA-Based Peak Detection

The Mu2e experiment at Fermilab imposes stringent requirements on the elimination of out-of-time beam in its pulsed proton beam - a requirement known as "extinction". We present a method to measure the out-of-time particle rates to calculate the level of extinction in the inter-pulse gaps. The proposed method utilizes an array of quartz Cherenkov radiators and photomultiplier tubes to detect particles scattered from a vacuum chamber in the M4 transfer beamline at Fermilab.The measurement will employ a new μTCA-based FPGA system for data acquisition and signal processing, utilizing real-time peak detection algorithms to count scattered beam particles. By integrating data over many transfers, the time profile of the out-of-time beam will be resolved to fractional levels relative to that of the in-time beam. These results are compared with G4beamline simulations to validate models of beam transport, dynamics, and extinction, providing critical input for optimizing beam delivery to Mu2e.

Hensley, Ryan [UC, Davis]

Dynamic CCS-EJ-SJ Database and Web Application - What's New

At the 2024 FECM/NETL Carbon Management Research Project Review Meeting, within the Carbon Transport and Storage Breakout Session 3, the presentation "Dynamic CCS-EJ-SJ Database and Web Application - What's New" highlights the critical tool designed to integrate environmental and social justice considerations into Carbon Capture and Storage (CCS) projects. Key features include an interactive dashboard for data access and visualization, which supports stakeholders in making informed decisions regarding CCS implementation, and updated data layers. The latest version enhances data integration and usability, providing a comprehensive resource for assessing the social and environmental impacts of CCS projects. There are 7 categories in the CCS EJSJ v2 database (released 03/31/2024): environmental justice, energy justice, economic justice, social justice, ecosystem assets, clean energy, and infrastructure. Most of the layers within each category have been updated in this version. As compared to the old database, there are 3 new categories in the v2 database: ecosystem assets, clean energy, and infrastructure.

Sharma, Maneesh

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie

Cybersecurity Standards for Distributed Energy Resources: Gaps and Harmonization Strategy

This report examines cybersecurity standards for Distributed Energy Resources (DERs) in light of their rapid growth and increasing integration into energy systems. It identifies critical gaps in existing frameworks, including inadequate coverage of DER-specific challenges, complexities in implementing comprehensive standards, integration issues with legacy systems, adoption hurdles for newer standards, and a lack of harmonization across regulatory landscapes. The analysis highlights vulnerabilities such as data integrity risks, unauthorized device control, and denial-of-service attacks across various DER technologies like solar PV, wind turbines, energy storage systems, and hydrogen fuel cells. The report proposes a harmonization strategy to address these deficiencies by developing unified cybersecurity requirements, certification programs, and training resources while fostering collaboration among stakeholders such as government agencies, industry groups, DER operators, manufacturers, and research institutions. A phased roadmap is outlined to refine and implement these measures through pilot testing and widespread adoption. Ultimately, the report underscores the urgent need for coordinated efforts to enhance DER cybersecurity and ensure the reliable operation of future energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY

MLSPICE: Machine Learning based SPICE Modeling Platform for Power Magnetics

Electrical power converters are critical to a wide range of applications ranging from renewable integration to transportation electrification, and can be a key factor determining the size, weight, and efficiency of energy conversion systems. Magnetic components are typically the largest and least efficient components in power electronics. While there have been major strides in the modeling and analysis of power semiconductor devices and circuit simulations, the necessary advances in the design of power magnetics have lagged. In this project, we have transformed the modeling and design of power magnetics with machine learning enabled methods and catalyze simultaneous disruptive improvements for ML-based power electronics design tools. A fully automated open-source machine learning based magnetics modeling platform – the MagNet project - with innovations in full stack have been developed to greatly accelerate the design process and provide new insights to magnetic material and geometry design. The ARPA-E funded MagNet platform contains three major building blocks: 1) a ML-Integrated Data Acquisition System (MIDAS): a highly automated data acquisition testbed which is capable of measuring a large number of magnetic cores with a wide range of electrical circuit excitations; 2) a ML-integrated Core Loss Model (MICLM): a machine-learning trained modeling method for modeling the core loss and saturation effects of magnetic materials for arbitrary excitation waveforms; 3) ML-guided Magnetics SPICE Simulation Tool (PMSPICE): a fully integrated CAD tool which can simulate the magnetics in SPICE. It can help the designers to quickly model the linear and non-linear characteristics of magnetic components and evaluate their behavior in SPICE simulations. The developed MagNet system has fully demonstrated the proposed performance target and has been open sourced to the entire power electronics community to advance the modeling and design of power magnetics from many different angles.

36 MATERIALS SCIENCE

A case study in contrastive learning information combination: Application to technical forensics of additive manufacturing filament source identification

Combination of information from disparate data sources into a single decision is a core challenge in many fields, including the field of technical forensics. Technical forensics (TF) utilizes technical characterization of questioned samples to determine properties of that sample; these properties are then used to infer information of forensic interest, such as provenance, age, or attribution. TF is utilized in traditional forensic applications, such as the attribution of material fragments from an explosive, and in nuclear forensic applications, such as the attribution of actinides which have been interdicted out of regulatory control. The challenge of combining information from disparate sources, described alternately by many terms including “Data Fusion” and “Data Integration”, is exacerbated in the technical forensics domain due to at least two factors: the challenge of interpreting each information source singularly, and the relatively small data set sizes available. Extensive literature exists attempting to combine technical forensics information sources, both in manual and automated processes. These attempts are often bespoke to the specific information sources (such as the bi-, tri-, or quad-isotope chart (Moody, Grant, and Hutcheon 2005)), with some emerging examples of simple early- and late- fusion (, respectively). Simultaneous to the information combination efforts described in the previous paragraph, the field of natural language processing attempted (and largely succeeded) in combining information from multiple non-technical information sources. The ecosystem of “multi-modal” language models, which can take text and images as input, and generate text and images as output, became large and diverse by 2025 (Khan et al. 2025). In a generalized sense, many of these methods are trained by learning neural networks which can convert raw text or images into a vector of numbers describing the text or image, hereafter called “embeddings” and the neural networks performing the conversion are called “embedders”. By using a separate embedder for text and images, finding coincident text and images (such as images with their captions), and optimizing the parameters of the embedders such that the embeddings for the text and the image are similar, the field has found a bridge between text and images (Girdhar et al. 2023). It is the contention of the authors of this report that this insight is not limited to text and images but instead can be extended to any modality which can be found coincidently. The subject of the rest of this report is the application of this method to example multi-modal technical forensic data. Some details about the data used in this report are not appropriate for this report, and are included in a companion report (PNNL-38669).

36 MATERIALS SCIENCE

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris

LYNM-PE1 Seismic Parameters from Borehole Log, Laboratory, and Tabletop Measurements

The goal of this work is to provide a database of quality-checked seismic parameters that can be integrated with the Geologic Framework Model (GFM) for the LYNM-PE1 (Low Yield Nuclear Monitoring – Physical Experiment 1) testbed. We integrated data from geophysical borehole logs, tabletop measurements on collected core, and laboratory measurements. We reviewed for internal consistency among each measurement type, documented the caveats of measurement conditions, and integrated lithologic logs to check the validity of outlier values. The resulting consolidated parameter tables can be used as inputs for modeling and analysis codes and are designed to interface with the GFM, which is being actively developed.

58 GEOSCIENCES

FL‐ADS: Federated learning anomaly detection system for distributed energy resource networks

Abstract With the ongoing development of Distributed Energy Resources (DER) communication networks, the imperative for strong cybersecurity and data privacy safeguards is increasingly evident. DER networks, which rely on protocols such as Distributed Network Protocol 3 and Modbus, are susceptible to cyberattacks such as data integrity breaches and denial of service due to their inherent security vulnerabilities. This paper introduces an innovative Federated Learning (FL)‐based anomaly detection system designed to enhance the security of DER networks while preserving data privacy. Our models leverage Vertical and Horizontal Federated Learning to enable collaborative learning while preserving data privacy, exchanging only non‐sensitive information, such as model parameters, and maintaining the privacy of DER clients' raw data. The effectiveness of the models is demonstrated through its evaluation on datasets representative of real‐world DER scenarios, showcasing significant improvements in accuracy and F1‐score across all clients compared to the traditional baseline model. Additionally, this work demonstrates a consistent reduction in loss function over multiple FL rounds, further validating its efficacy and offering a robust solution that balances effective anomaly detection with stringent data privacy needs.

Purohit, Shaurya [Iowa State University Ames Iowa

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics

AmpSuite

Seismic amplitudes offer vital information about explosion source characteristics, including discrimination and yield estimation. To take advantage of this, we developed an interactive Python package to measure, control data quality, generate broad area propagation models and perform discrimination and estimate yield. Propagation models are essential in support of transportable yield and broad area discrimination. The key benefit of this package will be its ability to continuously integrate data and new techniques. The AmpSuite framework will provide standardized, repeatable, and accurate model generation and characterization routines. The capability is crucial for monitoring agencies tasked with rapid and high-quality seismic event characterization. The AmpSuite software includes a series of independent modules to perform: • Direct Phase Amplitude Measurement and Storage • Coda Envelope Measurement and Storage • Data Quality Control • New Propagation Model Developments • Seismic Discrimination and Analysis • Yield Estimation and supporting utility software. The AmpSuite software provides comprehensive solutions for monitoring agencies seeking to optimize model generation and event analysis within a contemporary Python framework. Stakeholders (AFTAC) have begun to move towards the Python language for scientific analysis as a new workforce emerges.

Alfaro, Richard