Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Use of AI for Interpreting Technical Specifications for Power Uprates in Nuclear Power Plants

Powerpoint presentation. Background information provided on power plant uprates. Discussion of the current and proposed approaches to power plant uprates. Explanation of what data is used to draft a LAR. Methods such as retrieval augmented generation (RAG) and fine-tuning are discussed. Use case analysis is performed. Different failure types are examined. Conclusions are drawn from the analysis. Future work is proposed.

97 - MATHEMATICS AND COMPUTING↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES↗

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

DEVELOPMENT AND APPLICATION OF RISK ANALYSIS TOOLKIT FOR PLANT RESOURCE OPTIMIZATION

This paper presents the development of methods and tools that are being designed to optimize plant operations (e.g., maintenance/replacement schedules and optimal maintenance postures for plant components) in a manner that is more cost effective than current approaches and makes better use of available component health and cost data. These methods include both data- and model-based optimization methods. Model-based optimization methods directly include reliability and cost models to determine an optimal plant operational strategy. We consider gradient-based and evolutionary (based on genetic algorithms) optimization methods. The second class of methods target more specific use cases (e.g., project schedule optimization) and are not based on reliability models directly, but they require specific component reliability and cost data. This class of methods is based on variants of the knapsack problem with an aim to determine an optimal project schedule that maximizes the overall NPV. This paper also presents multi-objective methods designed to identify an optimal maintenance posture based on a Pareto frontier analysis. Rather than dictating the “right” tradeoff (i.e., identify the absolute best posture), we show how it is possible to perform a trade space exploration approach (i.e., identify value and costs of several postures and let the analysis account for desired value and cost metrics). This is performed by identifying maintenance postures that maximize value (e.g., system availability) and minimize operational costs, i.e., the Pareto frontier in a value-cost trade space. For all these methods we present detailed applicative examples that show their validity from a decision-making perspective.

97 - MATHEMATICS AND COMPUTING↗

Source Levels of In‐Cloud Air in Shallow Cumulus: Consistency Between Paluch Diagram and Lagrangian Particle Tracking

Abstract The Paluch diagram is a widely used tool for interpreting aircraft measurements of shallow cumulus clouds. A prior study conducted by Heus et al. (2008,https://doi.org/10.1175/2008jas2572.1) concluded that the source levels of in‐cloud air inferred from the Paluch diagram exhibit biases, sometimes of several hundred meters, in comparison to those derived from Lagrangian particle tracking. In this short study we revisit this comparison. The results indicate that the upper source levels of in‐cloud air determined from the Lagrangian Particle Tracking and the Paluch diagram are consistent, and the choice of statistical methods is crucial. The significance of this research lies in confirming the reliability of the Paluch analysis, enabling its confident application to aircraft data.

Meteorology & Atmospheric Sciences↗

Towards revealing intrinsic vortex-core states in Fe-based superconductors through statistical discovery

Abstract In type-II superconductors, electronic states within magnetic vortices hold crucial information about the paring mechanism and can reveal non-trivial topology. While scanning tunneling microscopy/spectroscopy (STM/S) is a powerful tool for imaging superconducting vortices, it is challenging to isolate the intrinsic electronic properties from extrinsic effects like subsurface defects and disorders. Here we combine STM/STS with basic machine learning to develop a method for screening out the vortices pinned by embedded disorder in iron-based superconductors. Through a principal component analysis of large STS data within vortices, we find that the vortex-core states in Ba(Fe 0.96 Ni 0.04 ) 2 As 2 start to split into two categories at certain magnetic field strengths, reflecting vortices with and without pinning by subsurface defects or disorders. Our machine-learning analysis provides an unbiased approach to reveal intrinsic vortex-core states in novel superconductors and shed light on ongoing puzzles in the possible emergence of a Majorana zero mode.

Guo, Yueming↗

Computing with a Chemical Reservoir

Contemporary computation is expensive, with large language models and artificial intelligence becoming more common in daily life. However, high-performance computing is reaching the limits in speed and energy expenditure, and domain science requires ever-increasing computational capacity, with simulations and data analysis pipelines ever-growing in complexity. As we progress towards post-exascale computation, with the associated high energy costs, new methods of energy-conscious computation are required. Novel analog and hybrid digital-analog systems can overcome these challenges, and chemical reactions offer a promising avenue. Computers based on chemistry can provide compact desktop devices with immense computational power. These devices are readily scalable by considering greater reaction systems or vessels, meeting the high-performance requirements for scientific workflows. In this article, we present ChemComp, a compilation pipeline for the conversion of ordinary differential equations into implementable chemical reactions. We then demonstrate the solving capabilities of ChemComp by emulating a potential chemical reservoir device. We leverage the multi-layer intermediate representation (MLIR) compiler framework to implement an expressive chemical reaction abstraction and propose a path for chemical reaction networks (CRNs) to represent mathematical problems effectively. Combined, we demonstrate a potential workflow that can harness chemistry’s computing power to create energy-efficient, high-performance computation systems for contemporary computing needs.

artificial intelligence↗

PNNL-Predictive-Phenomics/ProteoMeter

ProteoMeter is a Python package that assists in the statistical analysis of global proteomics, protein post-translation modification (PTM), and limited proteolysis (LiP) data. It contains batch correction, normalization, and statistical testing methods, as well as functions that "roll up" peptide-level data to the single-site level. It has a robust user configuration system, allowing it to flexibly integrate different types of experiment designs. For basic usage, a simple configuration file provides the essential functionality. Advanced users have access to the entire statistical pipeline for fine-tuning analyses. Processed data is easily exported to many common spreadsheet and data-frame formats.

Rozum, Jordan [Pacific Northwest National Lab]↗

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions↗

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Methods for Quantitative Thermal Analysis of Lithium Solid-State and Beyond Battery Safety

The use of differential scanning calorimetry (DSC) to measure the thermal behavior of individual components and electrolyte/electrode combinations is common. However, here we focus on DSC tests on an anode, cathode, and electrolyte (ACE) component combination over a temperature range that includes many of the phase transitions and key reactions (i.e., to 500 °C) that contribute to thermal runaway. This method can help quantify the complex reaction network in a full cell, thereby informing potential safety issues. Here, we used DSC heat flow data from a solid-state Li 0.43 CoO 2 +C+PVDF | LLZO | Li metal ACE sample and its components to quantify key factors affecting results. We focused on three areas: (1) ACE sample preparation and assembly in DSC pans, (2) DSC measurement parameters, and (3) heat flow analysis. Key points include the choice of component ratios (e.g., commercially relevant N:P capacity ratio), the importance of conductive carbon and binder, type of pan used, DSC ramp rate, and integration method used when dealing with broad and overlapping exothermic peaks. This work deepens the scientific basis and best practices for obtaining heat flow data from ACE samples for early-stage evaluation of solid-state and beyond battery safety.

25 ENERGY STORAGE↗

First observations of solar halo gamma rays over a full solar cycle

We analyze 15 years of Fermi-LAT data and produce a detailed model of the Sun’s inverse-Compton scattering emission (solar halo), which is powered by interactions between ambient cosmic-ray electrons and positrons with sunlight. By developing a novel analysis method to analyze moving sources, we robustly detect the solar halo at energies between 31.6 MeV and 100 GeV, and angular extensions up to 45° from the Sun, providing new insight into spatial regions where there are no direct measurements of the Galactic cosmic-ray flux. The large statistical significance of our signal allows us to subdivide the data and provide the first 𝛾-ray probes into the time variation and azimuthal asymmetry of the solar modulation potential, finding time-dependent changes in solar modulation both parallel and perpendicular to the ecliptic plane. Our results are consistent with (but with independent uncertainties from) local cosmic-ray measurements, unlocking new probes into astrophysical processes near the solar surface.

79 ASTRONOMY AND ASTROPHYSICS↗

In-Cell Deployment and First Use of Digital Image Correlation for In-Situ Strain Analysis of Irradiated Nuclear Fuel Rods During LOCA Transient

Digital image correlation (DIC) is a noncontact, optical method increasingly used across industries and research environments for acquiring multidimensional strain data. At Oak Ridge National Laboratory’s (ORNL’s) Severe Accident Test Station (SATS), DIC has been extensively applied to study the thermomechanical response of nuclear fuel claddings, yielding fundamental insights into material behavior. However, these efforts have focused exclusively on unirradiated materials and relied on the SATS out-of-cell infrastructure. Efforts over the past two years have been made to extend these capabilities to ORNL’s Irradiated Fuels Examination Laboratory SATS system within the hot-cells to enable testing of irradiated fuel cladding materials. Integrating DIC into the hot-cell SATS infrastructure presents unique challenges, including enabling remote operation of optical equipment and adapting auxiliary hot-cell systems for DIC implementation. This report discusses design, stand-up testing, and application of DIC to an irradiated nuclear fuel cladding segment. A DIC testing rig was successfully built and validated out-of-cell through extensive surrogate tests and calibrations. The rig and specialized DIC furnace were installed in the hot-cell, and a DIC test was successfully conducted. Results from that test correspond to expectations for Zr-based alloys and showed similar uncertainties compared to out-of-cell tests.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Field expedient stool collection methods for gut microbiome analysis in deployed military environments

ABSTRACT Field expedient devices and protocols for the collection, storage, and shipment of stool samples in deployed settings are needed for the advancement of microbiome research in military health. Relevant assessments include the evaluation of microbiome signatures associated with susceptibility to travelers’ diarrhea and recovery of gut function following infection. However, inherent biases in microbial measurements due to preservatives and sampling methods are unclear and should be assessed for an accurate evaluation of the microbiome. We performed shotgun metagenomic sequencing and compared the microbiome composition in paired fecal samples collected using Flinters Technology Associates (FTA) cards and OMNIgene (OG) Gut tubes, prior to and during international travel, from 49 adult participants, 39 of whom remained asymptomatic and 10 experienced travelers’ diarrhea. Higher concentrations of nucleic acid and sequencing libraries were observed in OG samples. A majority of genera (82.9%) were detected with both methods, and detections of genera limited to one collection method were not highly prevalent across samples and were present in extremely low relative abundances (<0.01%). Differences in beta diversity were largely explained by inter-individuality of microbiome composition, followed by the effect of collection method and timepoint-disease states. Differential abundance analysis indicated that Corynebacterium and Blautia were consistently higher in abundance across all groups with FTA and OG collection, respectively. The observed differences in microbiome composition between methods suggest the need for consistent and standardized protocols within a study. Overall, the data presented here could help guide the future design of fecal microbiome study protocols in field and military deployment settings. IMPORTANCE The assessment of field-deployable methods for fecal sample collection and storage is required to reliably capture samples collected in remote and austere locations. This study describes a comparative metagenomics analysis between samples collected by two different commercially available methods in a military-deployed setting. The results presented here are foundational for the future design of fecal microbiome study protocols in an operational context.

field study↗

Maximizing machine learning interatomic potential transferability for the discovery of the novel stellated octadecagon Bi18-Pt24 cage structure

Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributedStochasticNeighborEmbedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential’s high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 Å. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18⁢Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.

Vangheluwe, Raphaël [Université Paris-Saclay, CNRS↗

2016 U.S. Petroleum Fuels Life Cycle Baseline

The National Energy Technology Laboratory (NETL) has performed a well-to-wheels life cycle assessment (LCA) of petroleum production and refining for the United States (US), both from international and domestic oil sources. This analysis largely follows the same methods and framework established by Cooney et al., which performed a greenhouse gas (GHG) LCA of crude products for the 2014 data year (Cooney et al., 2017). Results are presented for six major petroleum products (gasoline, diesel, jet fuel, fuel oil, coke, bunker/residual fuel oil) across each of the five Petroleum Administration for Defense Districts (PADDs) and at the national US level. However, there are at least seven other refinery outputs (liquified petroleum gas, refinery fuel gas, hydrogen, petrochemical feedstocks, asphalt, sulfur) that are modeled but not shown in this report for brevity. Petroleum product amounts are compared against those reported by US Energy Information Administration (EIA) for each region.

02 PETROLEUM↗