Search NASASearch

SEARCH · Search NASA

Results for “Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH

What Is Data Analysis?

A quick guide to understand your data and use it to tell compelling stories. Data analysis helps develop insights for research projects, planning interventions, or systematic information gathering. This guide highlights important aspects of the data analysis process.

29 ENERGY PLANNING, POLICY, AND ECONOMY

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE

Adapt: A Weather Radar Data Analysis and Nowcasting Platform for Informed Adaptive Scanning

SF-26-021 Adapt is a data processing platform for real-time data analysis, short term prediction of targets convective cells and tracking for archived data. It provides tools for downloading, processing, segmenting, projecting, analyzing, and visualizing storm cell data from weather radar. The pipeline includes cell detection, motion estimation using optical flow, cell property extraction, and persistence to NetCDF and SQLite/Parquet for guiding adaptive scanning.

Raut, Bhupendra Ashokrao [Argonne National Laborat

Active multi-mode data analysis to improve fault diagnosis in AHUs

Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

HDG-1 Fiber Bragg grating data analysis

The main goal of the High dose graphite 1 Advanced test reactor experiment was to study nuclear grade graphite at high fluences. Additional supplementary optical fiber instrumentation was added to this long duration experiment for instrumentation development purposes. The supplementary instrumentation consisted of two pure silica core, fluorine doped cladding optical fibers each etched with 9 fiber Bragg gratings, one fiber being heat treated for 9 hours at 750 C and 16 hours at 750 C, the other being heat treated for 24 hours at 550 C and 48 hours at 650 C. Fiber Bragg gratings are known to have issues of measurement drift when in high temperature and high radiation environments like what is encountered in the Advanced test reactor. At the culmination of this experiment, the optical fibers saw ~1.3E21 n/cm2 total fluence, which is at the highest fluences that fiber Bragg gratings have been studied to date. Reported here is the analysis of this data including radiation induced shift and changes in sensitivity.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING

MapsTorch : automatic differentiation for X-ray fluorescence data analysis

X-ray fluorescence (XRF) is a popular spectroscopy technique for elemental analysis. Spectrum fitting and parameter tuning are at the core of XRF analysis and are conventionally manually intensive, especially for synchrotron experiments involving large amounts of diverse samples. This work introduces the automatic differentiation (AD) technique to XRF and an open-source package called MapsTorch. By transforming an analytical model of the XRF spectrum into a differentiable computation graph with AD, MapsTorch enables robust optimization of parameters and elemental intensities. We evaluate MapsTorch by conducting computational experiments on a large number of historical synchrotron XRF datasets and compare its performance with the currently practiced fitting tool NLopt. The results show that MapsTorch consistently achieves high-quality fits and often leads to better fitting quality than NLopt, particularly in tasks such as initial spectrum fitting and elemental intensity refinement. The robust performance of MapsTorch paves the way for developing automated and high-throughput XRF data analysis workflows to handle the increasing data volumes expected from next-generation synchrotron facilities.

X-ray fluorescence

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)

Evaluation of Machine Learning Models for Automated Data Analysis in In-Service Nuclear Power Plant Inspections

The commercial nuclear power industry is facing a potential shortage of certified nondestructive evaluation (NDE) analysts to meet future in-service inspection demands. Automated data analysis (ADA) currently supports human inspectors in tasks such as eddy current evaluations for steam generator examinations. Machine learning (ML) systems are nearing the capability to pass performance demonstration tests for ultrasonic testing (UT) inspections of reactor pressure vessel upper head penetrations in nuclear power plants (NPPs). Current research and development is focused on assisted analysis (AA) of ADA versus fully automated examinations. This presentation will cover assessment of ML flaw detection on dissimilar metal weld (DMW) piping joints.

36 MATERIALS SCIENCE

Doppler Backscattering Data Analysis and Integrated Modeling with OMFIT

One Modeling Framework for Integrated Tasks (OMFIT) is a widely used software tool in the magnetic fusion research community. OMFIT provides magnetic fusion energy researchers with a framework for the development of special-purpose physics modules. This paper describes an OMFIT physics module pertaining to the Doppler Backscattering (DBS) fusion plasma diagnostic. DBS measures density fluctuations and flow velocity through plasma scattering of electromagnetic waves. The OMFIT DBS module was developed to analyze experimental DBS data and facilitate modeling of DBS systems installed on multiple tokamak devices. The OMFIT DBS module is designed to support several analysis workflows: detailed analysis of experimental data, experimental planning, and theory-based synthetic diagnostic modeling. The DBS module uses integrated modeling by leveraging other OMFIT physics modules to perform tasks related to DBS, e.g. ray/beam–tracing simulations, edge-localized mode–synchronized data analysis, magnetic equilibrium reconstruction, and fitting kinetic profile data. Furthermore, this paper describes several supported workflows and serves a reference for the OMFIT DBS module.

Doppler backscattering

Data for 3-Hydroxypropionic Acid Recovery from Fermentation Broth through Novel Downstream Processing: Technoeconomic Analysis

This study develops and validates a simplified, fully solvent-free downstream processing (DSP) strategy for high-purity recovery of 3-hydroxypropionic acid (3-HP) from real fermentation broth containing 62.3 g/L of 3-HP. Optimized activated carbon treatment achieved 98% color removal, while Amberlite IRA-67 was operated at pH 4.5 and 30 °C to minimize product loss. This is the first integrated demonstration of a fully solvent-free DSP enabling recovery of bio-based 3-HP as both a solid sodium salt and a concentrated aqueous solution, supported by techno-economic analysis. At lab scale, the process achieved 77.3% recovery of sodium 3-HP with 83.2% (w/w) purity and produced a 30% (w/v) aqueous solution. Techno-economic analysis yielded minimum selling prices of $0.551/kg for the solution and $0.892/kg for the salt, both below target thresholds for cost-competitive bio-acrylic acid production. Overall, these results demonstrate an efficient, scalable, and economically viable industrial pathway for 3-HP recovery.

Bioproducts

Brazilian CBP - Technoeconomic analysis data

This data is related to the paper entitled "Techno-economic analysis of sugarcane bagasse and straw conversion into cellulosic ethanol via consolidated bioprocessing". That features the evaluation of sugarcane bagasse and straw conversion to ethanol at stand-alone facilities generating electricity from residues. The following scenarios were evaluated: Conventional, featuring hydrothermal pretreatment, fungal cellulase, and yeast fermentation (current commercial standard); Mid-term consolidated bioprocessing (CBP), relying on bagasse solubilization without pretreatment or cotreatment; and Mature CBP, incorporating cotreatment but no pretreatment and considering significant technological advance of the CBP. Available here are the spreadsheets used for Material and Energy balance calculation, Capital and Operational costs estimation and Cash flow analysis. Also available are the description and python code used for Monte Carlo analysis of the ethanol and capital investment variations. This data can be used as a source to implement other techno-economic analysis in the biorefinary context.

09 BIOMASS FUELS

Establishing Data Analysis Pipeline for Bulk ATAC-Seq Datasets

We developed an analysis pipeline for transposase-accessible chromatin sequencing (ATAC-Seq) data derived from bulk samples, which brings together publicly available R packages in addition to command-line tools designed for analysis of bulk ATAC-Seq data and can be run on any computer running a Linux-like operating system such as Ubuntu or Apple OSX.

97 MATHEMATICS AND COMPUTING

AGC-2 Graphite Preirradiation Data Analysis Report

This report describes the specimen loading order and documents all preirradiation examination material property measurement data for graphite specimens contained within the Second Advanced Graphite Capsule (AGC 2) irradiation capsule. The AGC 2 capsule is the second in six planned irradiation capsules comprising the Advanced Graphite Creep (AGC) test series. The AGC test series is used to irradiate graphite specimens in order to garner quantitative data necessary for predicting the irradiation behavior and operating performance of new nuclear grade graphites. This testing will ascertain the in service behavior of the graphite for pebble bed and prismatic very high temperature reactor designs. Similar to the First Advanced Graphite Capsule (AGC 1) preirradiation examination report, material property tests were conducted on specimens from 18 nuclear grade graphite types. However, AGC 2 tested an increased number of specimens (i.e., 512) prior to loading them into the AGC 2 irradiation assembly. All AGC 2 specimen testing was conducted at Idaho National Laboratory from July 2009 to August 2010. This report also details the specimen loading methodology for graphite specimens inside the AGC 2 irradiation capsule. The AGC 2 capsule design requires “matched pair” creep specimens that have similar dose levels above and below the neutron flux profile mid plane. This provides similar specimens with and without an applied load. Analysis in this document utilizes the neutron flux profile calculated for the AGC 2 capsule design, the capsule dimensions, and the size (i.e., length) of the selected graphite specimens to create a stacking order that produces “matched pairs” of graphite specimens above and below the AGC 2 capsule elevation mid point, thus providing specimens with similar neutron dose levels.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]

Operating Experience Data Analysis for Digital Instrumentation and Control System Reliability and Risk Assessment in Nuclear Power Plants

The implementation of advanced digital instrumentation and control (DI&C) systems in U.S. nuclear power plants (NPPs) can bring significant advancements in reliability, monitoring, and control capabilities. However, these systems also introduce new challenges, particularly in assessing risks such as common-cause failures (CCFs) and establishing robust reliability estimates for DI&C components. Addressing these challenges is critical for ensuring the safe and efficient operation of NPPs. Recently, Idaho National Laboratory was tasked by the U.S. Nuclear Regulatory Commission (NRC) to conduct a DI&C reliability study using operating experience data from the nuclear industry. The two operating experience data sources for the study are the Institute of Nuclear Power Operations’ Industry Reporting and Information System (IRIS) and the NRC’s Licensee Event Report database which is hosted at Idaho National Laboratory at https://lersearch.inl.gov/LERSearchCriteria.aspx. This report provides a comprehensive examination of DI&C systems, including their architecture, operational advantages, and associated challenges. It reviews existing industry DI&C studies and failure mode taxonomies, along with reliability data from various industries. Through a detailed analysis of these databases, the study provides insights into DI&C system performance. Considerations should be given to incorporate DI&C failure data into the NRC's Integrated Data Collection and Coding System and updating the Reliability and Availability Data System to support ongoing DI&C reliability studies. Recommendations are also provided for modeling DI&C reliability and CCF in probabilistic risk assessment, thereby supporting risk-informed decision-making and enhancing the reliability and safety of NPPs.

22 GENERAL STUDIES OF NUCLEAR REACTORS

In Situ Data Analysis Through Physics-informed Tensor Decompositions (LDRD Final Report)

We introduce a new low-dimensional model of high-dimensional numerical simulation data based on low-rank tensor decompositions. Our new model aims to minimize differences between the model data and simulation data as well as functions of the model data and functions of the simulation data. This novel approach to dimensionality reduction of simulation data provides a means of directly incorporating quantities of interests and invariants associated with conservation principles associated with the simulation data into the low-dimensional model, thus enabling more accurate analysis of the simulation without requiring access to the full set of high-dimensional data. Computational results of applying this approach to two standard low-rank tensor decompositions of data arising from simulation of combustion and plasma physics are presented.

97 MATHEMATICS AND COMPUTING