Search NASASearch

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber

Absorption Correction for Reliable Pair Distribution Functions from Low Energy X-ray Sources

This paper explores the development and testing of a simple absorption correction model for processing powder X-ray diffraction data from Debye−Scherrer geometry laboratory X-ray experiments. This may be used as a preprocessing step before using PDFGETX3 to obtain reliable pair distribution functions (PDFs). Various experimental and theoretical methods for estimating μR were explored, and the most appropriate μR values for correction were identified for different capillary diameters and X-ray beam sizes. We identify operational ranges of μR where a reasonable signal-to-noise ratio is possible after correction. A user-friendly software package, DIFFPY.LABPDFPROC, is presented that can help estimate μR and perform absorption corrections with a rapid calculation for efficient processing.

Absorption

Quantifying Impacts of Biomass Pelletization on Fast Pyrolysis Using a Single-Particle Reactor, X-ray Computed Tomography, and Computational Modeling

The pore structure and density of lignocellulosic feedstocks dictate intraparticle transport phenomena and thereby play an important role in thermochemical conversion processes such as fast pyrolysis for biofuel and biochemical production. Variations in microstructure are inherent from different biomass species and can be introduced by preprocessing techniques such as cutting and pelletization. Morphological changes also occur during conversion and lead to vastly different pore structures and behavior during pyrolysis, which impact required conversion times and product distributions. The current work presents a comprehensive comparison of fast pyrolysis of neat and pelletized pine feedstocks, which includes single-particle experiments, modeling, and 3D imaging by X-ray computed tomography (XCT). The particle-scale model included anisotropic heat and mass transport in a shrinking particle with pyrolysis reactions based on the CRECK mechanism with boundary conditions informed by reactor-scale simulations of the single-particle reactor. The models were validated by measurements of the temperature and mass loss from single-particle pyrolysis experiments of neat and pelletized pine. Quantitative analysis of XCT geometries revealed that pyrolytic conversion yielded chars with increased porosity and permeability compared to the unpyrolyzed materials, along with decreased tortuosity and anisotropy. Pelletization of the pine feedstock resulted in a much denser, less permeable material, which converted slower and produced more residual char after pyrolysis compared to neat pine. The results from particle modeling revealed that accounting for the dynamic and anisotropic heat and mass transport caused by differences in pore structure is critical to achieving agreement with experimental results. Overall, this study highlights the dramatic differences in conversion behavior imparted by pelletization and the importance of capturing microstructural attributes in computational models to guide the design and optimization of pyrolysis processes for specific biomass feedstocks.

09 BIOMASS FUELS

Process Feasibility Analysis of Waste Biomass Valorization to Biochar and Bio-Oil via Slow and Fast Pyrolysis

The United States has abundant biomass and waste feedstock to support the nation's energy addition and affordability targets. Pyrolysis, a thermochemical conversion process, decomposes lignocellulosic feedstocks into liquid, solid, and gaseous fuels that can contribute to the domestic production of biofuels, biopower, and bioproducts. Growing private sector interest in this technology is a key motivation for this comprehensive techno-economic process modeling analysis of a respective biorefinery that includes feedstock preprocessing, slow and fast pyrolysis, and product separation to bio-oil, biochar, and syngas hydrocarbons. Results show that biochar from slow pyrolysis could achieve minimum selling prices (MSPs) of $\$$188-$\$$260/t, competitive with reported market values, while bio-oil from fast pyrolysis is estimated to yield MSPs of $\$$6.49-$\$$9.68/GGE, approximately twice conventional fuel benchmarks. Sensitivity analysis identifies feedstock cost, product yield, and scale as primary cost drivers, while scenarios involving biochar carbon credits and high value applications may substantially improve economics. Overall, these results suggest that continued innovation in feedstock logistics, process integration, and market development will be critical to achieving economically viable and scalable bioproducts.

09 BIOMASS FUELS

Aggregation Methods for Quantifying PTM and Structural Changes in Bottom-Up Proteomics

Bottom-up proteomic workflows rely on sequential preprocessing steps, commonly including peptide-to-protein aggregation (“roll-up”), to enhance data reliability and interpretability. While roll-up is effective for protein-centered analyses, it may be suboptimal for applications focused on post-translational modifications (PTMs) or protein structural changes, such as limited proteolysis–mass spectrometry (LiP-MS). Here, we investigate how different roll-up strategies influence site-level quantification in PTM differential analysis. Moreover, we introduce a novel site-centric roll-up approach tailored for LiP-MS, which quantifies proteolytic fragments rather than solely tryptic peptides. We benchmark these methods through simulation studies, comparing their sensitivity and specificity in detecting structural and PTM-driven changes. We found that the median and mean roll-up methods outperform the sum method in both PTM and LiP proteomics, and site-level quantification in LiP outperforms peptide-level quantification. Our findings offer the first systematic, data-driven guidance for selecting roll-up techniques in site-level proteomic analyses, with implications for both PTM-focused and structural proteomics studies.

aggregation

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE

Explainable tokamak-agnostic forecasting of fusion plasma instability via megahertz turbulent fluctuations

Scientific applications of artificial intelligence (AI) often remain limited by device-specific training and unexplained “black-box” approaches, creating fundamental barriers to cross-system generalization. This challenge is critical for nuclear fusion, where future reactors will have limited operational data for AI training. Here, we demonstrate that our neural network, trained solely on megahertz-scale turbulence measurements from one machine (DIII-D), forecasts Type-I edge localized mode (ELM) onsets in a different tokamak (KSTAR) through zero-shot weight transfer following physics-consistent preprocessing without device-specific retraining. Through an explainable AI framework combining gradient-weighted class activation mapping with physics validation, we reveal that our network can internalize physics relationships governing the ELM instabilities rather than memorizing device-specific patterns. The network perceives spatiotemporal features that correlate consistently with independently calculated instability growth rates, magnetohydrodynamic stability limits, and pedestal structure dynamics. Statistical analyses of dimensionally-reduced saliency features reveal the identical triangular features between the saliency representations, instability growth rates, and prediction probability across tokamaks, providing evidence that our forecasting system can show tokamak-agnostic generalization. This work contributes to a foundation for explainable scientific AI systems, where cross-system developments are essential for transcending traditional domain-specific constraints.

AI

ChemPren: a new and economical technology for conversion of waste plastics to light olefins

With the ever-increasing demand for plastics, sustainable recycling methods are key necessities. Here, the current plastics industry can manage to recycle only 10% of the 400 million metric tons of plastic produced globally. Waste plastics, in the current infrastructure, land up mostly in landfills. Although a lot of research efforts have been spent on processing and recycling co-mingled mixed plastics, energy-efficient sustainable and scalable routes for plastic upcycling are still lacking. Catalytic valorization of waste plastic feedstock is one of the potential scalable routes for plastic upcycling. Silica-alumina based materials, and zeolites have shown a lot of promise. A major interest lies in restricting catalyst deactivation, and refining product selectivity and yield for such catalytic processes. This article highlights ChemPren technology as a clean energy solution to waste plastic recycling. Co-mingled, mixed plastic feedstock along with spray dried, attrition resistant, ZSM-5 containing catalysts is preprocessed with an extruder to form optimally sized particles and fed into a fluidized bed reactor for short contact times to produce selectively and in high yields ethylenes, propylenes and butylenes. This techno-economic perspective indicates that the ChemPren technology can produce propylene at $\$$0.16 per lb, whereas the current selling price of virgin propylene is $0.54 per lb. This technology can serve as a platform for mixed plastic upcycling, with more advancements necessary in the form of robust and resilient catalysts and reactor operation strategies for tuning product selectivity.

25 - ENERGY STORAGE

Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments

We introduce a computational framework that integrates artificial intelligence (AI), machine learning, and high-performance computing to enable real-time steering of neutron scattering experiments using an edge-to-exascale workflow. Focusing on time-of-flight neutron event data at the Spallation Neutron Source, our approach combines temporal processing of four-dimensional neutron event data with predictive modeling for multidimensional crystallography. At the core of this workflow is the Temporal Fusion Transformer model, which provides voxel-level precision in predicting 3D neutron scattering patterns. The system incorporates edge computing for rapid data preprocessing and exascale computing via the Frontier supercomputer for large-scale AI model training, enabling adaptive, data-driven decisions during experiments. This framework optimizes neutron beam time, improves experimental accuracy, and lays the foundation for automation in neutron scattering. Although real-time experiment steering is still in the proof-of-concept stage, the demonstrated potential of this system offers a substantial reduction in data processing time from hours to minutes via distributed training, and significant improvements in model accuracy, setting the stage for widespread adoption across neutron scattering facilities and more efficient exploration of complex material systems.

97 MATHEMATICS AND COMPUTING

Forming a database to study reversed magnetic shear from the National Spherical Torus eXperiment using machine learning

Achieving a long-lived reversed magnetic shear (RMS) target plasma in the National Spherical Torus eXperiment Upgrade will require developing various sustainment scenarios. To help with the ongoing plasma control efforts, the development of a new analysis for the motional Stark effect (MSE) diagnostic using a machine learning algorithm, namely, MSE-ML, is described. MSE-ML will be used to identify patterns during RMS discharges, some of which suffer magnetohydrodynamic (MHD) events resulting in current redistribution and monotonic q-profiles. A database consisting of q and magnetic shear profiles is being constructed primarily based on the existing National Spherical Torus eXperiment data with equilibrium reconstructions constrained by the magnetic field pitch angle profile measured using the multi-channel MSE diagnostic. An unsupervised k-means clustering of the data is developed to study the RMS formation as a function of time. The initial clustering from the q-profiles shows significant differences in both amplitude and the duration of the RMS period. As a goal, the clustering results that detect and distinguish shots with substantial and sustained RMS are to be used as a preprocessing step in a supervised algorithm to identify the underlying conditions that lead to long-lasting improved confinement with RMS. Another aim of the MSE-ML study is to identify precursors of RMS-destroying MHD events in either derived data such as the q-profile or directly measured data such as the magnetic field pitch angle profile.

Uzun-Kaymak, I. U. (ORCID:0000000276251493)

Identification of Distorted Gamma-Ray Signature Patterns Using Digital Filtering and Auto-Associative Memory Implemented with a Hopfield Neural Network

The detection and identification of radioactive sources in search applications involve analyzing passive gamma-ray emissions from high-level radioactive materials. This process uses a mobile detector-spectrometer in a complex field test environment. Recently, the use of artificial intelligence for gamma-ray spectrum analysis has shown promising results. However, challenges persist in identifying isotopic signatures from spectral measurements that may be distorted due to source shielding, random variations in natural radioactive background, or insufficient measurement time to obtain clear spectral lines. Here, this paper presents a novel intelligent signature recognition method that combines digital filtering techniques with an artificial Hopfield Neural Network (HNN). The HNN leverages auto-associative memory to store training sample patterns and match them with incoming gamma spectra from distorted sources. It restores the testing sources’ measurements by finding the closest matching signature patterns in the spectral library. Before HNN recognition, the measured spectrum undergoes preprocessing with a digital image filter to reduce fluctuations. Performance of the proposed method is evaluated using a set of gamma-ray spectra measured with a sodium iodide detector. The data collected include measurements from six pure samples: 241 Am, 60 Co, 137 Cs, 192 Ir, 239 Pu, and 235 U, which are used for training and validation (i.e. six cases). Additionally, the data set contains 24 distorted synthesized sources with various fluctuating backgrounds. Test results demonstrate the potential of the proposed method to accurately recognize the correct isotope with high precision, achieving an accuracy rate exceeding 85%. Furthermore, the proposed method exhibits superior performance compared to the conventional multiple regression fitting and simple feedforward neural network methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Advancing the STS Neutron Moderator Design with an Automated Optimization Workflow and Unstructured Mesh Modeling

With the Second Target Station approaching its final design phase, a detailed neutronics evaluation of its critical components is necessary. Optimizing the dimensions of the two cold-source moderators that are at the heart of this facility presents a multi-objective optimization problem for which an accurate geometric description is crucial. We have applied a fully automated optimization workflow in which a detailed unstructured mesh geometry is automatically generated with Attila4MC, starting from a parametrized CREO geometry followed by preprocessing with SpaceClaim. With this geometry, a MCNP run is performed to calculate the brightness metrics, which are subsequently provided to the optimization algorithm in DAKOTA that provides new parameters and drives the optimization loop until convergence. In this paper, we show the results of the analysis that are used for the final design of the cylindrical and tube moderator. The optimization simulations provide a refinement to and confirmation of the conclusions of the previous design iteration. Additional to the optimization, a sensitivity study is performed to study the effect of minor geometry changes, which is important for the final engineering design. In conclusion, with these studies, we demonstrate that the automated workflow and high-fidelity unstructured mesh modeling are efficient tools for a thorough design evaluation.

DAKOTA

Point spread function deconvolution using a convolutional autoencoder

A major issue in optical astronomical image analysis is the combined effect of the instrument’s point spread function (PSF) and the atmospheric seeing that blurs images and changes their shape in a way that is band and time-of-observation dependent. In this work we present a very simple neural network based approach to nonblind image deconvolution that relies on feeding a convolutional autoencoder (CAE) input images that have been preprocessed by convolution with the corresponding PSF and its regularized inverse, a method which is both conceptually simple and computationally less intensive. We also present here, a new approach for dealing with limited input dynamic range of neural networks compared to the dynamic range present in astronomical images.

79 ASTRONOMY AND ASTROPHYSICS

Adaptively coupled phase retrieval in multi-peak Bragg coherent diffraction imaging

Recent advances in Bragg coherent diffraction imaging (BCDI) experimental techniques permit routine measurement of multiple Bragg peaks from a single crystalline grain. The resulting images contain the full lattice distortion vector field which can be differentiated to provide lattice strain and rotation. With the advent of fourth-generation synchrotron light sources, such multi-peak datasets are produced at high rates, facilitating the need for rapid phase retrieval of the multiple peaks and subsequent image analysis. Here we describe and demonstrate a new implementation of a coupled phase retrieval technique for multi-peak BCDI which simultaneously treats each Bragg peak of the dataset and produces a three-dimensional image of the crystal's morphology and lattice distortion field. In addition, this method uses the redundant information contained in the various Bragg diffraction patterns to detect and suppress spurious signal appearing on the detector in a subset of the measurements. Compared with manual data editing, adaptive coupling produces a more consistent phase profile in reciprocal space and sharper surfaces in direct space, with no significant difference in computational cost. These improvements reduce the need for manual preprocessing and enable robust high-throughput analysis of multi-peak BCDI data, supporting near-real-time strain microscopy at modern synchrotron facilities.

36 MATERIALS SCIENCE

Leveraging BERT and Network-Based Attention Analysis for Identifying Treatment Milestones in EHRs

This study introduces a sophisticated data-driven framework for analyzing Electronic Health Records (EHRs) using transformer-based models to identify and disentangle overlapping treatment contexts. The framework leverages a preprocessing pipeline that transforms structured procedural codes into semantically enriched descriptive text, enabling the use of attention mechanisms to cluster medical events into treatment milestones—cohesive and distinct components of care processes. The methodology is rigorously validated using synthetic datasets derived from the MIMIC-III database, designed to simulate the heterogeneity and overlapping procedural contexts characteristic of real-world EHR scenarios. Quantitative evaluation highlights the framework’s robustness in disentangling concurrent care pathways, with attention metrics and unsupervised clustering approaches demonstrating the ability to preserve intra-context relationships while distinguishing inter-context dependencies. By addressing challenges inherent in data heterogeneity, this approach provides a foundation for uncovering complex treatment patterns, advancing clinical decision-making, and optimizing resource allocation in diverse healthcare environments.

Kim, Minsu [ORNL] (ORCID:0000000224185535)

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING

Scalable Multi-Facility Workflows for Artificial Intelligence Applications in Climate Research

Earth observation satellites and earth system models are sources of vast, multi-modal datasets that are invaluable for advancing climate and environmental research. However, their scale and complexity pose significant challenges for processing and analysis. In this paper we discuss our experiences in developing and using a scientific research application using an automated multi-facility workflow that orchestrates data collection, preprocessing, artificial intelligence (AI) inferencing, and data movement across diverse computational resources, leveraging the Advanced Computing Ecosystem Testbed at the Oak Ridge Leadership Computing Facility (OLCF). We demonstrate that our workflow can be seamlessly integrated and orchestrated across research facilities managed by different federal agencies, thus allowing users to extract new scientific insights from climate datasets. The experimental results indicate that the multi-facility workflow significantly reduces processing time, enhances scalability, and maintains high efficiency across varying workloads. Notably, our workflow processes 12,000 high-resolution satellite images in just 44 seconds using 80 workers distributed across 10 nodes on the OLCF systems. Such high throughput is essential for dynamic tokenization and sharding of petascale satellite data for distributed AI model training and inferencing at scale across thousands of GPUs.

Kurihana, Takuya [ORNL] (ORCID:0000000156698565)