Search NASA⌕ Search

SEARCH · Search NASA

Results for “data provenance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

OPEN-Augmented Reality GUI for Bioenergy Crop Phenotyping and Precision Agriculture (Donald Danforth Plant Science Center Final Scientific Technical Report)

The project led by the Donald Danforth Plant Science Center, in collaboration with Arizona State University, George Washington University, and Saint Louis University, has made significant strides in advancing the phenotypic analysis of bioenergy crops through the development of an innovative AI processing pipeline. This initiative was primarily funded by ARPA-E, with additional cost-sharing provided by the participating institutions. The project successfully utilized a variety of sensors—3D scanners, thermal, RGB, and hyperspectral—to refine algorithms for data-driven trait signature identification and improve the classification and visualization of plant traits. The developed AI processing pipeline is capable of handling the complex, multidimensional data characteristic of dynamic agricultural environments. 1) Contributions to understanding: The research has advanced the field of plant phenomics by showcasing the synergistic use of various sensor data to enhance the precision of trait analysis in bioenergy crops. Through the integration of 3D scanners, thermal, RGB, and hyperspectral sensors, the project has developed robust data-driven trait signature algorithms and visualization techniques. These innovations have facilitated detailed monitoring and management of plant traits, providing vital insights into plant growth dynamics and stress responses. Further, the project has broadened our understanding of how machine learning can be effectively applied in multi-sensor environments to refine trait analysis. By leveraging diverse datasets, the research has not only improved the accuracy of phenotypic assessments but also established a versatile methodological framework that can be extended beyond agriculture to other fields requiring detailed phenotypic analysis. 2) Technical effectiveness and economic feasibility: The AI processing pipeline developed in this project demonstrated significant technical effectiveness, achieving high throughput analysis of extensive phenotypic data and meeting targeted accuracies. This system exemplified the capability of advanced machine learning technologies to efficiently manage and analyze large, complex datasets. Economically, the implementation of the project-developed pipelines may offer substantial cost savings across multiple sectors. It enhances data analysis processes and significantly reduces the need for manual data interpretation, thereby decreasing both the time and resources required. 3) Public benefit: The project has significantly broadened the scope of agricultural methodologies to enhance phenotypic analysis, with potential applications in various sectors beyond agriculture. Additionally, the initiative fostered an enriching educational and collaborative environment, significantly enhancing the technical skills of participants. It also made substantial contributions to the scientific community by providing open-access data sets and tools, encouraging ongoing research and development across various disciplines. Overall, the project not only met its scientific goals but also showcased the extensive utility of integrating advanced machine learning and sensor data analysis technologies. These advancements have proven instrumental in driving forward both theoretical research and practical applications, setting a strong foundation for future explorations and innovations in data-driven science.

60 APPLIED LIFE SCIENCES↗

Consist v0.1.0

A Python library for provenance tracking, intelligent caching, and data virtualization in scientific simulation workflows. It automatically records code, configuration, and input data to skip redundant computations and enables querying results across many runs without manual bookkeeping. Designed to support multi-model simulation workflows like the BEAM CORE toolset at LBL, but designed to be extensible to a wide range of research workflows. Combines lineage tracking features as provided by OpenLineage with deterministic hashing like SnakeMake, and adds powerful analysis tools on model outputs.

Needell, Zachary [Lawrence Berkeley National Labor↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING↗

Alabama Carbon Storage: Bringing Data to the People

The Gulf Coastal Plain of Alabama has proven potential for geologic carbon storage and current interest in the area for large carbon capture and storage (CCS) projects is high. Extensive CCS relevant data exist in the records of the Geological Survey of Alabama and State Oil and Gas Board of Alabama, however, most of this data is not publicly available or is scattered in separate databases, file cabinets, and tables in publications. The “Alabama Carbon Storage: Data Sharing and Engagement” (ACS-DSE) project seeks to accelerate the responsible development of large CCS projects in the Gulf Coastal Plain of Alabama and offshore in state waters through a publicly accessible database of geologic carbon storage models and data across the region. The ACS-DSE draws on the over 150 years of geologic research and over 20 years of experience in CCS research to place relevant geologic, geophysical, and infrastructure data on a single web platform. Datasets available will include formation depths and elevations, geologic structures, reservoir properties, digital well logs (LAS files), existing penetrations, and geologic models. In addition to downloadable datasets, links to CCS related regulatory agencies and other sources of information will be included (for example, Class VI UIC permitting regulations and pipeline regulations). By making these datasets and models available in commonly used formats on a public website, the project will increase transparency in decision making and decrease the data acquisition time for industry.

01 COAL, LIGNITE, AND PEAT↗

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

Feedforward-feedback ammonia control at a water resource recovery facility based on a digital twin with hybrid model

Ammonia-based aeration control (ABAC) at full-scale Water Resource Recovery Facilities (WRRFs) can be challenged by diurnal loading and transport delays. This work addressed these challenges using a hybrid feedforward–feedback controller built on Activated Sludge Model 1 (ASM1), marking the first full-scale deployment to pair a mechanistic feedforward core with data-driven corrections. The objectives were to improve ammonia setpoint tracking, assess performance of the mechanistic model when enhanced with data-driven corrections, and document full-scale operation. The hybrid model incorporates two data-driven components: (1) a Mechanistic Error Forecasting Engine (MEFE), consisting of a multivariate linear regressor and a long short-term memory (LSTM) ensemble. Defying expectations, low-parameter models outperformed more complex alternatives, reducing the mechanistic error by 71%. (2) A Residual Oscillation Forecasting Engine (ROFE), based on Fast Fourier Transform, reduced the remaining error by another 35%. Two proportional–integral (PI) feedback loops further (i) trim the feedforward output and (ii) eliminate residual controller error in the final aerobic zone. In full-scale operation, the controller reduced mean-squared error (MSE) by 94% over the baseline and produced more stable dissolved oxygen (DO) setpoints. Overall, it was proven that layering multi-timescale data-driven models on a mechanistic core can yield reliable ABAC performance at WRRFs.

54 ENVIRONMENTAL SCIENCES↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Rhenium Isotope Reconnaissance of Uranium Ore Concentrates

Exploration of natural isotopic variations of the element rhenium (Re) is in its infancy, with initial studies revealing isotopic fractionation in a variety of geological materials. Here, in this work, we investigate Re isotope variation as a new geochemical tool, given its redox-sensitive properties and affinity for organic matter and sulfides. In this work, Re abundance and isotope ratio data were collected from uranium ore concentrates (UOCs) across a variety of depositional ages, locations, geologic settings, and deposit types. Ore types from which the UOC were derived include sandstone, unconformity, and quartz-pebble (QP) conglomerate. To isolate Re from the U-rich matrix of UOCs, a new purification method utilizing DGA ion exchange resin was developed. We found that UOCs exhibit a wide range of Re isotope ratios, with sandstone ore-derived UOCs having the isotopically lightest values, QP conglomerate ore-derived UOCs having the heaviest, and unconformity ore-derived UOCs in between (with some overlap with sandstone UOCs). The Re isotope ratio range observed in UOCs extends previously reported values by more than a factor of two. Industrial processing (e.g., incomplete recovery of Re from ore, contamination, fractionation during processing) may play a role in the isotopic variability in the UOCs. However, systematic differences between ore types suggest that the depositional setting is a significant factor. For nuclear forensic investigations, Re isotopic compositions combined with data from other isotopic systems provide geochemical signatures that can aid in provenance assessment of UOCs. Regardless of the specific causes for the wide range of Re isotope ratios in UOCs, these initial data indicate Re is a promising tool for nuclear forensic investigations on samples from early in the nuclear fuel cycle.

58 GEOSCIENCES↗

Real Time-Optimal Power Flow-Based Distributed Energy Resource Management System (DERMS)

This project aims to promote lab-proven clean energy technology to commercially scalable versions of the technology, integrate the technology with broader systems, provide extended performance data, and validate the manufacturability and reliability of the technology. The lab-proven technology, RT-OPF DERMS, was developed and validated through previous U.S. Department of Energy-funded efforts, including Advanced Research Projects Agency-Energy funding under the Network Optimized Distributed Energy Systems program and Holy-Cross Energy High Impact Project. In the Advanced Research Projects Agency-Energy Network Optimized Distributed Energy Systems project, the RT-OPF DERMS was developed and implemented in multiple hardware platforms, demonstrating its performance and capabilities in the lab and field environments. The technology was also evaluated and matured via a participation in the U.S. Department of Energy I-Corps program, whose goal is to pair teams of researchers with industry mentors for an intensive 2-month training in which the researchers define technology value propositions, conduct customer discovery interviews, and develop viable market pathways for their technologies. These activities indicate the high technology maturity and Technology Readiness Level of the RT-OPF DERMS.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Opening doors to physical sample tracking and attribution in Earth and environmental sciences

Physical samples and their associated data and metadata underpin scientific discoveries across disciplines and can enable new science when appropriately archived. However, there are significant gaps in current practices and infrastructure that prevent accurate provenance tracking, reproducibility, and attribution. For most samples, descriptive metadata are often sparse, inaccessible, or absent. Samples and associated data and metadata may also be scattered across numerous physical collections, data repositories, laboratories, data files, and papers with no clear linkage or provenance tracking as new information is generated over time. The Earth Science Information Partners (ESIP) Physical Samples Curation Cluster has therefore developed guidance for scientific authors on ‘Publishing Open Research Using Physical Samples.’ This involved synthesizing existing practices, gathering community feedback, and assessing real-world examples. We identified improvements needed to enable authors to efficiently cite and link Earth science samples and related data, and track their use. Our goal is to help improve discoverability, interoperability, and reuse of physical samples, and associated data and metadata. Though primarily focused on the needs of Earth and environmental sciences, these guidelines are broadly applicable.

58 GEOSCIENCES↗

exfor_client

A lightweight Python client and CLI for interacting with the [EXFOR Web API](https://nds.iaea.org/exfor/x4guide/API/). This tool enables searching, retrieving, and parsing experimental nuclear data — including uncertainties, covariance information, and metadata — while preserving provenance.

Grosskopf, Mike [Los Alamos National Laboratory]↗

Particle Filter Based Inference Testing

The primary intent of PAR-FIT (Particle Filter based Inference Testing) is to provide hard inductive evidence that a machine learning model is capable and proven for an individual test input. By examining training data used to form the underlying model functional correlation, an estimate of the reliability that a model will make the correct prediction can be made. The Sequential Probability Ratio Test is used to derive a qualitative evaluation for reliability based on hypothesis testing. The PAR-FIT framework achieves this by implementing a particle filter and the sequential probability ratio test algorithms on the machine learning model training data to determine relevancy of new individual test samples to the training dataset. The kernel function evaluates the local proximity and density of training data used to derive a prediction outcome. Particles are used to probabilistically determine which training data to evaluate for proximity. For test samples that are within a close proximity to and surrounded by multiple training data points, the evaluated reliability of the prediction is high. For test samples that are anomalies not represented by the training dataset, in low density data clusters, or are far from existing data points, the evaluated reliability is low as insufficient training evidence exists to suggest the model is capable of making the correct prediction. Sequential Probability Ratio Test is further used to determine when a hypothesis on whether a signal can be rejected or accepted for use. The ratio test collects sequence information from the particle filter to test whether the signal is anomalous or normal via hypothesis testing of the underlying distributions.

Chen, Edward [Idaho National Laboratory (INL), Ida↗

New K-feldspar Pb isotope results for Mesozoic arc crust in the Pacific Northwest, U.S.A. and Canada: comparison with the Mojave-Salinia province of southern California and Implications for Baja-BC

Measurements of lead isotopic compositions in detrital K-feldspar have been increasingly used as a tool to assess sediment provenance. We compiled a database of previously published Pb isotope data from 700 bedrock K-feldspar samples and 1,423 age-corrected bedrock whole rock samples from western North American igneous and metamorphic bodies. Additionally, we report 66 new K-feldspar Pb isotope data for plutons throughout the Pacific Northwest region of the United States and British Columbia. Results show that the Pb isotope values of plutonic K-feldspar depend on the isotopically juvenile or evolved nature of underlying crust. Samples obtained from the mid Cretaceous – mid Eocene Coast Plutonic Complex, North Cascades, and Intermontane superterrane that occur west of the initial 87 Sr/ 86 Sr (Sri) = 0.706 isopleth exhibit a highly restricted 207 Pb/ 206 Pb and 208 Pb/ 206 Pb values centred upon 0.83 and 2.03, respectively. Conversely, rocks overlying older continental crust further east such as the Middle Jurassic – Late Cretaceous Omineca crystalline belt, Idaho batholith, and Boulder Batholith exhibit far greater variation of Pb isotope values that parallel the 100 Ma isochron calculated from a two-stage Pb evolution model. We demonstrate that Pb isotopic results from the Idaho and Boulder Batholith region can be used to define distinctive subregions for Pb isotopic provenance analysis, and compare these signatures to the Mojave-Salinian batholith of southern California and western Arizona, as these two areas have previously been proposed as source regions for extraregional sediment that was deposited within the Nanaimo Basin during the Campanian – Maastrichtian. Future Pb isotopic analysis of detrital K-feldspar from the Nanaimo Basin of southwestern British Columbia may effectively distinguish between potential extraregional sources separated by thousands of kilometres.

Coast Plutonic Complex↗

Hardware-in-the-Loop Testing of Wide-Area Damping Controller for Field Implementation in Large-scale Power Grid

In our previous work, an adaptive measurement-driven wide-area damping controller (WADC) for suppressing inter-area oscillations has been proposed and a hardware prototype was developed and validated through hardware-in-the-loop tests. As a continuation of the work, this paper introduces a WADC software prototype to handle the realistic challenges for field implementation in the control room of the power grid. The WADC software is developed and operated as an openPDC adapter with a graphical user interface (GUI) to monitor the WADC inputs and output, the communication delays and other variables. The software prototype has been fully tested through an enhanced hardware-in-the-loop (HIL) test setup. Its performance is verified under various realistic communication uncertainties, such as random time delays and data losses, with different communication protocols. The experiment results have proven the WADC software can deliver sufficient damping to suppress the targeted oscillation mode in handling various communication uncertainties for future field deployment.

Jia, Xinlan↗

ATcT — Active Thermochemical Tables Python Interface

SF-25-140 atct is a lightweight, Python client for the ATcT v1 API that enables programmatic access to high-accuracy thermochemical data and turnkey reaction-enthalpy analysis. The package implements full v1 endpoint coverage (species lookup by ATcT ID, name, formula, SMILES, InChI, CAS RN; covariance queries; health checks) with robust error handling, retries, and environment-based configuration for local/production endpoints. Beyond data retrieval, atct provides rigorously implemented reaction calculators that propagate uncertainties via either (i) a conventional independent-errors method (0 K or 298.15 K) or (ii) covariance-aware propagation using provided covariances at 298.15 K. Typed data classes ensure transparent, reproducible data structures and carry ATcT Thermochemical Network (TN) version identifiers for provenance. Dual import paths and comprehensive examples facilitate integration into research pipelines, enabling reproducible thermochemical calculations, automated validation, and downstream method development.

Bross, DavidHamilton [Argonne National Laboratory ↗

Optimizing Facility Operations by Applying Machine Learning to the Army Reserve Enterprise Building Control System (Final Report)

Thousands of U.S. Department of Defense (DoD) buildings have building automation systems (BASs) and/or advanced meters. Although these systems have a wealth of data, performance optimization requires time and expertise to review and act on that information. Machine learning (ML) can provide automated and actionable insights to controls operators. This demonstration implemented proven ML methods on the Army Reserve Enterprise Building Control System. ML refers to algorithms that “learn” from data and improve their performance on a given task over time. In the buildings domain these tasks range from predicting future energy consumption, to identifying operational issues before faults occur, to optimizing control decisions. To learn, ML requires input data, which – for buildings – typically consists of instrument data such as energy consumption data and subsystem controls information such as set-point temperatures, and context data consisting of information such as the physical location of the building, the area of the building, and the weather. ML models use the relationships learned from the input data to make predictions with new, previously unseen, data. The team was able to investigate and successfully implement the following ML use cases: labeling consumption data as anomalous or non-anomalous; baseline whole-building load prediction (unknown fault status); fault detection (validation not possible); and site prioritization for energy-related projects. Due to the constraints of the project, interventions were not able to be implemented during the demonstration; therefore, assessments of operational cost savings and maintenance avoided could not be performed. The project has been presented at two leading national building conferences and two additional publications to peer-reviewed journals are currently in preparation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predicting High‐Resolution Spatial and Spectral Features in Mass Spectrometry Imaging with Machine Learning and Multimodal Data Fusion

Recent advancements in molecular Mass Spectrometry Imaging have sparked interest in integrating high spatial resolution methods with molecular mass-spectrometry-based chemical imaging. Fusion-based algorithms have proven effective in generating high spatial-resolution molecular mass spectra. However, a significant challenge stems from the differing physical mechanisms underlying image generation and data upsampling techniques, potentially leading to discrepancies in integrated information channels. Integrating physical constraints into data processing workflows is essential to tackle this issue. In this study, we propose an innovative approach that merges data from Fourier transform ion cyclotron resonance (FTICR), time-of-flight matrix-assisted laser desorption/ionization, and time-of-flight secondary ion mass spectrometry imaging techniques. By leveraging FT-ICR's unparalleled spectral resolution and ToF-SIMS's exceptional spatial resolution, we achieve submicron spatial resolution, enabling the observation of intact molecular species with remarkable spectral precision. Canonical correlation analysis is employed to incorporate physical constraints. Through sophisticated image processing and machine learning techniques, the results of this fusion hold significant promise for advancing our comprehension of complex systems and unveiling concealed molecular intricacies.

canonical correlation analysis↗