Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Condition-Based Maintenance of a Circulating Water System of a Canadian Nuclear Power Plant using Machine Learning and Statistical Tools

Canada Deuterium Uranium pressurized-heavy-water reactors (PHWR) are a type of nuclear power plant that generate clean and reliable energy. The scope of this work is to automate data analysis methodologies to inform a condition-based maintenance strategy of a circulating water system (CWS) of a PHWR. The multiunit CWS provides a continuous supply of water to cool steam condensers, even during transient scenarios, thereby improving the thermal efficiency. This work aims to develop a machine learning (ML) based approach to detect anomalies in heterogeneous data of a CWS in a PHWR to help inform a predictive maintenance strategy. The heterogeneous data include textual and numeric time series data for a PHWR. Natural-language-processing (NLP)-based models are used to analyze textual data contained in work orders and operator logs and an event-timeseries correlation detection method is applied to assist anomalies diagnoses for CWS. An ML model Robust Linear Model (RLM) is also used to remove the seasonal variations in the system variable distributions based on distributions of environmental variables. A machine learning model, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), trained on both original data and data without any seasonal variations will then be used to detect if an anomaly exists. Thus, by moving to an automated methodology to detect, classify, and forecast anomalies, the maintenance strategy would be based on component condition instead of a time-based schedule.

97 - MATHEMATICS AND COMPUTING↗

buhito

buhito is a Python library for graph analysis and machine learning. Graphs can represent networks with objects as nodes and their relationships as edges. buhito focuses on graphlet methods that study graphs through enumerating their component subgraphs to enable interpretable and fast models of complex systems. The package provides tools for different algorithmic designs for computing, analyzing, and applying graphlets to research problems such as machine learning, data compression, and anomaly detection in graph-structured data. A central feature is performing decomposition data analysis on graphs for machine learning models. Implemented in Python and built upon open-source scientific libraries such as NetworkX, NumPy, and SciPy, buhito provides high-performance methods for researchers exploring the mathematical and computational foundations of graphlet analysis applicable to systems of different sizes.

Pimonova, Yulia↗

TEAMER: Twin Ocean Power Wave Energy Converter Comprehensive Overview

These files collectively provide a comprehensive overview of the testing process, data analysis, and validation for the Twin Ocean Power device tested at the O.H. Hinsdale Wave Research Laboratory, supported by TEAMER funding. This resource includes an overview of power results for a series of 7 trials. The files included in this comprehensive overview include a comprehensive log sheet for each trial, a summary of all trials, and processing scripts for the raw data. It includes all raw data in .tsv and MATLAB compatible formats, an average power chart, angular velocity charts for each trial, trial metrics, and power output files. This resource includes images of the Twin Ocean Power Wave Energy Converter device components and movement during testing and video recordings of each trial.

16 TIDAL AND WAVE POWER↗

Lithium-Ion Battery Design for Grid-Scale Energy Storage App

A software that delivers parameters from energy storage system (ESS) to container, rack, module and single cell design, as well as data analysis on arbitrage energy and frequency regulation of ESS in different regions, has been developed. The Lithium-ion Battery Design for Grid-scale Energy Storage App V1.0 has the capability to output the system, module and cell design with the energy, power, capacity, group method, cost of single cell, and single cell test protocol which break down from input energy storage system data in different regions. The default chemistry of the battery is LiFePO4 and graphite. The energy density of the graphite/LiFePO 4 pouch cell ranges from 100 Wh/kg to 200 Wh/Kg in the software. Graphite/LiFePO 4 pouch cell (up to 1Ah in lab) manufacturing line is also built and can be used to evaluate the test protocol, moreover, for electrolyte evaluation in other ESMI seedling projects. The software enables rapid prototyping to accelerate energy storage research, development, and manufacturing.

Liu, Dianying [Pacific Northwest National Laborato↗

Wilkins: HPC in situ workflows made easy

In situ approaches can accelerate the pace of scientific discoveries by allowing scientists to perform data analysis at simulation time. Current in situ workflow systems, however, face challenges in handling the growing complexity and diverse computational requirements of scientific tasks. In this work, we present Wilkins, an in situ workflow system that is designed for ease-of-use while providing scalable and efficient execution of workflow tasks. Wilkins provides a flexible workflow description interface, employs a high-performance data transport layer based on HDF5, and supports tasks with disparate data rates by providing a flow control mechanism. Wilkins seamlessly couples scientific tasks that already use HDF5, without requiring task code modifications. We demonstrate the above features using both synthetic benchmarks and two science use cases in materials science and cosmology.

HPC↗

EGS Collab Experiment 2: Microseismic Monitoring

This dataset contains continuous seismic waveform data recorded during stimulation and thermal circulation tests for the Enhanced Geothermal Systems (EGS) Collab Experiment #2, conducted from February to September 2022 at the Sanford Underground Research Facility in Lead, South Dakota. This experiment aimed to study and validate models of geothermal systems by injecting high-pressure fluids into rock formations 1200-1500 meters below the surface, inducing microseismic events. The seismic monitoring system included 16 three-component accelerometers and a 24-channel hydrophone array, installed in boreholes surrounding the test area. Data were recorded at high sampling rates using a continuous waveform recording system to monitor seismic activity in real time. The dataset contains the raw data stored in binary format, with files named based on timestamps, and includes calibration certificates for some sensors to facilitate corrections to real units. Users are strongly advised to consult the accompanying detailed report, which outlines the experimental setup, sensor specifications, installation procedures, and data processing methods. The report also describes important nuances, such as the hardware filters on hydrophones, sensor calibration details, and the naming conventions for the recorded data. Proper use of this dataset may require familiarity with seismic data analysis tools, such as the Obspy Python package, and an understanding of the SEED naming conventions used for channel identification.

15 GEOTHERMAL ENERGY↗

CalTestBed - Delphire - Testing and Evaluation of Delphire Sentinel System (CRADA Final Report)

The Delphire Sentinel is a modular fire detection and communications system operating as a mobile field unit, with low voltage DC power supplied by onboard photovoltaics (PV) and batteries. The Sentinel addresses several aspects of fire detection, communications and data analysis. The Sentinel's mobility enables it to be rapidly deployed and operate independently of existing power and communications networks. The duration of independent operation depends critically on the energy consumption of the systems and performance of the onboard PV and battery. The purpose of this testing is to ascertain the power draw and energy consumption of the Delphire Sentinel prototype system under several operational states, including various data transfer packet sizes, transmission time and frequencies, and communication pathways (Wi-Fi, cellular, satellite) expected to be encountered in field deployments. It will also include procedures to test the ability of the Sentinel to operate for extended periods without loss of functionality. Based on results from energy and power measurements, and anticipated duty cycles in field deployments, we will model annual system autonomy (e.g. loss of load probability) for off-grid operation in representative locations.

47 OTHER INSTRUMENTATION↗

Nondestructive In Operando Imaging of Thin Film Composite Membrane Compaction Enhanced by AI-Based Segmentation

Reverse osmosis (RO) membranes are essential for desalination and water reuse, yet their permeability declines in high-pressure applications due to membrane compaction. This study investigates the structural and functional responses of commercial brackish, seawater, and high-pressure RO membranes at applied pressures up to 120 bar using a multiscale, nondestructive in operando scanning electron microscopy (iSEM) imaging platform. The iSEM technique reveals progressive densification across the composite membrane structure, which correlates with observed declines in water and solute permeance. To quantify these structural changes with greater fidelity, we combined X-ray computed tomography with AI-based segmentation enabling precise analysis of pore size distribution and thickness of the polysulfone support layer. Compared to traditional thresholding, AI segmentation accurately delineates material phases and void spaces, enhancing the reproducibility and resolution of morphological assessments. The results demonstrate that compaction-induced reductions in porosity and thickness strongly impact membrane transport properties. These findings provide mechanistic insights into the compaction behavior of RO membranes and underscore the potential for advanced imaging and AI-driven data analysis to guide the design of next-generation membranes with improved mechanical resilience and operational longevity.

13 HYDRO ENERGY↗

Exponential concentration in quantum kernel methods

Kernel methods in Quantum Machine Learning (QML) have recently gained significant attention as a potential candidate for achieving a quantum advantage in data analysis. Among other attractive properties, when training a kernel-based model one is guaranteed to find the optimal model’s parameters due to the convexity of the training landscape. However, this is based on the assumption that the quantum kernel can be efficiently obtained from quantum hardware. In this work we study the performance of quantum kernel models from the perspective of the resources needed to accurately estimate kernel values. We show that, under certain conditions, values of quantum kernels over different input data can be exponentially concentrated (in the number of qubits) towards some fixed value. Thus on training with a polynomial number of measurements, one ends up with a trivial model where the predictions on unseen inputs are independent of the input data. We identify four sources that can lead to concentration including expressivity of data embedding, global measurements, entanglement and noise. For each source, an associated concentration bound of quantum kernels is analytically derived. Lastly, we show that when dealing with classical data, training a parametrized data embedding with a kernel alignment method is also susceptible to exponential concentration. Our results are verified through numerical simulations for several QML tasks. Altogether, we provide guidelines indicating that certain features should be avoided to ensure the efficient evaluation of quantum kernels and so the performance of quantum kernel methods.

97 MATHEMATICS AND COMPUTING↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Measurement of Stark-split beam and carbon charge exchange emissions for simultaneous B-field and temperature/rotation analysis at DIII-D

A set of two newly designed, single-channel Czerny–Turner spectrometers has been deployed at the DIII-D tokamak for measurements of the motional Stark effect (MSE) split beam emission and the C6+ (CVI) carbon charge exchange recombination (CER) emission at high spectral (δλ = 0.13 nm) and temporal (1–5 kHz) resolution. High throughput optics (f/# = 2.8) allow for good signal-to-noise at high time resolution using fast EMCCD detectors. The MSE emission allows for spectral fitting of the magnitude and direction of the local B-field, while the carbon emission yields local ion temperature and toroidal rotation information. To reduce so-called Doppler broadening of the MSE emission, a new channel-specific variable lens-masking approach has been developed. Experimental data collected from the 2023 DIII-D experimental campaign demonstrate the signal quality and instrument fidelity for both diagnostic measurements. Moreover, initial CER data analysis shows a clear evolution of the toroidal rotation during edge localized modes. Initial progress on the advanced MSE model, including a new validated ray-trace model of the DIII-D collection optics, is shown via sensitivity analysis.

Instruments & Instrumentation↗

Characterizing Interaction Uncertainty in Human-Machine Teams

With the increasing use and adoption of artificial intelligence (AI), the reliability of modern data systems will be driven by a tighter teaming between human experts and intelligent machine teammates. As in the case of human-human teams, the success of human-machine teams will also rely on clear communication about mutual goals and actions. In this paper, we combine related literature from cognitive psychology, human-machine teaming, uncertainty in data analysis, and multi-agent systems to propose a new form of uncertainty: interaction uncertainty for characterizing bidirectional communication in human-machine teams. We map the causes and effects of interaction uncertainty and outline potential ways to mitigate uncertainty for mutual trust in a high-consequence real-world scenario.

uncertainty, data analytics, interaction, trust, h↗

Comparative Analysis of DNA LLM Classification Techniques Using Intra-Layer Feature Extraction with Autoencoder Stacks [Poster]

This project conducts a comparative analysis of DNA LLM classification techniques using Evo2, Grover, and UTRML, focusing on intra-layer feature extraction in Evo2. By extracting features from multiple layers of Evo2 and integrating them into an autoencoder stack with a binary classification head, we evaluate its effectiveness in classifying genomic sequences compared to smaller DNA language models. My findings demonstrate that Evo2 outperforms Grover and UTRML in classification accuracy on a dataset provided by department 08625, CAO2021, while UTRML offers competitive performance with lower computational costs. This study highlights the potential of advanced embedding techniques in enhancing genomic data analysis and informs future research in bioinformatics.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

The Single Event Error (SEE) test and analysis of the CMS Endcap Timing Layer readout chip

The ETROC2, the first full size and full functionality prototype chip for the CMS Endcap Timing Layer readout, is strategically designed to meet the SEE immunity requirements of detector operation with the low power constraint. The triplicated periphery and pixel I2C configuration registers are designed with self-correction feature. The pixel readout control is centralized in the global readout and fully triplicated. The pixel readout is not triplicated, instead protected with power-efficient one-bit correction Hamming code. The TMR protection of the on-pixel threshold calibration can be turned off allowing the detection of the beam spot during the beam test by checking the bit-flips of the internal memory cells. In the initial proton beam test in January 2024, the chip readout process did not hang throughout the tests. The Hamming code correction strategy works because the error corrected TDC data were observed in the data frames. The error-injection simulation is performed to analy ze the small number of bit-flips in the configuration registers. We also performed SEE testing with a heavy ion beam in April and the data analysis is ongoing. The detailed design on the SEE immunity and the testing as well as simulation results will be presented, including follow-up SEE testing results in May and June 2024.

Gong, Datao↗

Analysis of the electrical double layer using electrochemical X-ray photoelectron spectroscopy

The element-sensitivity of X-ray spectroscopies offers the potential to disentangle the individual chemistries of water, ions, and adsorbates at the electrode-electrolyte interface in an element-by-element manner. However, targeted experimental design is needed to establish interface-sensitive in situ X-ray spectroscopy in a realistic electrochemical environment. Here, we demonstrate how electrochemical X-ray photoelectron spectroscopy (EC-XPS) in the dip-and-pull geometry can be used to specifically probe the behavior of ions in the electrical double layer. Taking the case study of a polycrystalline Au foil in 50 mM KClO 4 electrolyte, we tracked the electrochemical response of interfacial K + cations across a broad potential range. We show how, in combination with modeling, key parameters such as the potential of zero charge (PZC), ion packing behavior, dielectric saturation, and the electrostatic potential decay in the double layer can be extracted from the data. Importantly, we also analyze how the experimental conditions and non-idealities can influence the results and put forward criteria for reliable experimentation and data analysis.

Dip-and-pull↗

Using Best Basis Inventory Data to Direct Strategies for Real-Time Monitoring of Hanford High Level Waste

The proposed Direct Feed High Level Waste (DFHLW) approach for processing high-level tank waste at Hanford is intended to reduce processing time by bypassing the Pretreatment Facility and transferring waste directly from the tank farm to the WTP HLW vitrification facility. This processing strategy could reduce or eliminate the washing and leaching steps that would have occurred in the Pretreatment facility. Operation of the vitrification facility is subject to chemical and radiological limits protecting safety (e.g. Waste Acceptance Criteria, or WACs) and process quality (e.g. Process Control Limits, or PCLs). Without washing and leaching, there is a greater risk of exceeding the WACs and PCLs. Hanford process engineers have devised blending strategies based on known chemical and radiological composition, volumes, and solids loadings of individual layers within each waste tank. These blending campaigns succeed in predicting a processing strategy that does not exceed the WACs and PCLs. However, the calculations do not ascribe uncertainties to the tank analysis data, quantities of material taken from the tanks to make the blend, or potential for mixing of layers within tanks. In order to confirm that a process strategy is working, it would be advantageous to have inline or at-line analytical instrumentation installed in the processing facilities that deliver measurement results in real time.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

ASCR Workshop Position Paper: Challenges and Opportunities in High Energy Physics

High energy particle physics and cosmology concern themselves with estimating fundamental parameters of nature, such as the masses and interactions of fundamental particles like the Higgs boson and the rate of expansion of the universe. In doing so, they analyze exabyte-scale datasets, some of the largest in all of science, and face many challenges in subsequent data analysis. These challenges are shared between the two disciplines, but we focus on particle physics to highlight one specific domain. In particle physics, the standard method for estimating parameters involves performing Monte Carlo (MC) integration as a function of both parameters of interest and nuisance parameters using an expensive simulator, counting the number of observed collision events (i.i.d. samples) from an experiment in the corresponding integration domains, and forming a Poisson likelihood function. This likelihood function is then used in a Frequentist manner to construct a maximum likelihood point estimate (MLE) and confidence set for the parameters. To sufficiently populate the high-dimensional integration domains, simulators consume billions of CPU-hours annually and produce hundreds of petabytes of intermediate output data. Several techniques have been developed to: optimize definitions of the integration domains so as to be maximally sensitive to a particular subset of parameters, efficiently estimate the integrals, and build robust surrogate models by interpolating between integral evaluations at different parameter points. One can view this whole endeavor as classical Simulation-Based Inference (SBI).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗