Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Automating Anomaly Detection for Target systems at Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory, produces the world’s most intense pulse neutrons beams. An accelerated proton beam is directed into a mercury target to generate neutrons via spallation. The target system accounted for over 40% of the overall downtime of the facility in 2022. Thus, early detection in anomalies in the target systems can enable taking corrective actions to avoid failures and reduce downtime. Fault prognostics and anomaly detection in accelerators, both at SNS and outside, has largely focused on the beam side. This paper presents one the first studies exploring leveraging machine learning to automate the detection of anomalies in the target system. The target system consists of over 30 different interconnected subsystems, and the present work focuses on the mercury process system as a use case. Analyzing data from 28 process variables from 2022 and 2023, tree-based and reconstruction-based algorithms are employed to detect anomalies in archived data. The algorithms detected previously unreported anomalies, several of which were deemed alert worthy by human experts, particularly those found by reconstruction-based algorithms. Using data from each production run in the accelerator increased the generalizability of the models in time. Efforts are now underway to implement a workflow for incorporating human feedback to update the models and evaluating performance on unseen data. The models will eventually be integrated into the existing System Tracking and Reliability system with a web interface for automated anomaly detection and reporting along with a pathway for incorporating human feedback for model updates.

Raj, Anant [ORNL] (ORCID:0000000306711244)↗

A co-registered in-situ and ex-situ dataset from wire arc additive manufacturing process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data-centric approach emphasizes leveraging sensor data available throughout the production process to optimize performance. Integration of extensive data analysis provides opportunities for improving precision, reducing waste, and enhancing the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes a comprehensive description of the deposition process, process parameters, welding characteristics and acoustic data collected in-situ, and X-Ray Computed Tomography data of the build.

42 ENGINEERING↗

Post-processed surface meteorological, air-sea flux, SST, wave, and ship navigation/position merged data from NOAA Ship Pisces

These are post-processed surface data from the NOAA Ship Pisces, merged from several instruments. They consist of meteorological, ocean, ship navigation/position, and surface air-sea flux quantities calculated at three different time frequencies. Because the averaging of directions can be complex, they are provided as a courtesy. R1 in the filename indicates "revision 1" or version 1. R2 will include surface waves (TBD).

17 WIND ENERGY↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Toward memory-efficient melt pool monitoring: a classification framework using event-based imaging and sparse sensing technique

Vision sensors like CMOS and CCD cameras are often used for in-process monitoring of melt pools in laser-based additive and welding processes, but they require transferring large amounts of data and computational processing resources. Event-based neuromorphic imagery, on the other hand, detects only the change in pixel intensity, thus potentially reducing the data amount and latency. With an event imager, this study develops a framework for melt pool condition classification, including image construction, time scale selection, optimal pixel selection, and sparse classification, to achieve a highly memory-efficient scheme. These are based on sparse sensing techniques with singular value decomposition (SVD) and QR pivoting, the two fundamental matrix transformations for linear dimensionality reduction. The framework is then validated by classifying a controlled experiment by exciting various mode shapes of liquid gallium pools of varying depths (3, 6, and 8 mm). At 200 pixels, the classifier can reach overall accuracy of 75%, while at 2000 pixels (0.013% of the total possible pixels), the accuracy is nearly 90% (89.86%). At the same number of pixels, random selection can only achieve 46% and 67%, respectively. The memory savings of the sparsely sampled event data compared to a conventional imager is about 500 times. In addition to performance, implementation and limitations of the framework are also discussed.

42 ENGINEERING↗

bmdrc: Python package for quantifying phenotypes from chemical exposures with benchmark dose modeling

Though chemical exposures are known to potentially have negative impacts on health, including contributing to chronic diseases such as cancer, the quantitative contribution of risk is not fully understood for every chemical. A commonly used approach to quantify levels of risk is to measure the proportion of organisms (such as a total number of zebrafish on a plate or mice in a cage) with abnormal behavioral responses or morphology at increasing concentrations of chemical exposure. A particular challenge with processing the proportional data from these assays is the appropriate estimation of chemical concentration levels that result in malformations or acute toxicity, as these values typically vary between experimental measurements. The recommended approach by the Environmental Protection Agency (EPA) is to fit benchmark dose curves with specific filters and model fitting steps, which are crucial to properly processing the proportional data. Several tools exist for the fitting of benchmark dose response curves, but none are standalone Python libraries built to process both morphological and behavioral data as proportions with all the EPA recommended filters, filter parameters, models, and model parameters. Thus, here we present the benchmark dose response curve (bmdrc) Python library, which was built to closely follow these EPA guidelines with helpful visualizations of filters and fitted model curves, and reports for reproducibility purposes. bmdrc is open-source and has demonstrated utility as a support package to an existing web portal for information on chemicals (https://srp.pnnl.gov). Our package will support any toxicology analysis where the response is a proportional value at increasing levels of a concentration of a chemical or chemical mixture.

Superfund↗

HERMES

HERMES - High-speed Event Retrieval and Management for Enhanced Spectral imaging code. This code is meant to unpack and process neutron imaging data from the TPX3Cam made by Amsterdam Scientific Instruments. The TPX3Cam utilizes the Timepix3 chip in a single photon counting mode of image acquisition. The photon counting data streaming off the TPX3Cam needs to be process and analyzed in order to create final images. Therefore we are developing both python and cpp codes to allow for users of the TPX3Cam to efficiently analyze data and create images.

Long, Alexander↗

Processed Soil Respiration at the TRACE experimental Warming project, Aug 2015 - Sep 2017, Sabana, Luquillo, Puerto Rico

This data package contains processed measurements of soil carbon dioxide (CO₂) efflux collected using LI-COR LI-8100 soil respiration chambers at the Tropical Responses to Altered Climate Experiment (TRACE) located at the Sabana Field Research Station near Luquillo, Puerto Rico. The TRACE site is a mature, closed-canopy tropical wet forest within the Luquillo Experimental Forest. These data quantify soil surface CO₂ fluxes from both ambient (control) and experimentally warmed plots to evaluate how long-term soil warming affects belowground carbon cycling in tropical ecosystems. The data files include time-series tables of CO₂ flux (µmol CO₂ m⁻² s⁻¹), soil temperature (°C), and ancillary environmental variables, stored in comma-separated values (CSV) format and viewable with any text editor, spreadsheet, or statistical software (e.g., R, Python, Excel). Associated metadata describe plot identifiers, measurement intervals, and processing steps. These data were generated to address the research question: How does sustained soil warming influence soil respiration and carbon flux dynamics in tropical wet forests?

54 ENVIRONMENTAL SCIENCES↗

High‐speed 4‐dimensional scanning transmission electron microscopy using compressive sensing techniques

Abstract Here we show that compressive sensing allows 4‐dimensional (4‐D) STEM data to be obtained and accurately reconstructed with both high‐speed and reduced electron fluence. The methodology needed to achieve these results compared to conventional 4‐D approaches requires only that a random subset of probe locations is acquired from the typical regular scanning grid, which immediately generates both higher speed and the lower fluence experimentally. We also consider downsampling of the detector, showing that oversampling is inherent within convergent beam electron diffraction (CBED) patterns and that detector downsampling does not reduce precision but allows faster experimental data acquisition. Analysis of an experimental atomic resolution yttrium silicide dataset shows that it is possible to recover over 25 dB peak signal‐to‐noise ratio in the recovered phase using 0.3% of the total data. Lay abstract : Four‐dimensional scanning transmission electron microscopy (4‐D STEM) is a powerful technique for characterizing complex nanoscale structures. In this method, a convergent beam electron diffraction pattern (CBED) is acquired at each probe location during the scan of the sample. This means that a 2‐dimensional signal is acquired at each 2‐D probe location, equating to a 4‐D dataset. Despite the recent development of fast direct electron detectors, some capable of 100kHz frame rates, the limiting factor for 4‐D STEM is acquisition times in the majority of cases, where cameras will typically operate on the order of 2kHz. This means that a raster scan containing 256^2 probe locations can take on the order of 30s, approximately 100‐1000 times longer than a conventional STEM imaging technique using monolithic radial detectors. As a result, 4‐D STEM acquisitions can be subject to adverse effects such as drift, beam damage, and sample contamination. Recent advances in computational imaging techniques for STEM have allowed for faster acquisition speeds by way of acquiring only a random subset of probe locations from the field of view. By doing this, the acquisition time is significantly reduced, in some cases by a factor of 10‐100 times. The acquired data is then processed to fill‐in or inpaint the missing data, taking advantage of the inherently low‐complex signals which can be linearly combined to recover the information. In this work, similar methods are demonstrated for the acquisition of 4‐D STEM data, where only a random subset of CBED patterns are acquired over the raster scan. We simulate the compressive sensing acquisition method for 4‐D STEM and present our findings for a variety of analysis techniques such as ptychography and differential phase contrast. Our results show that acquisition times can be significantly reduced on the order of 100‐300 times, therefore improving existing frame rates, as well as further reducing the electron fluence beyond just using a faster camera.

Robinson, Alex W.↗

Validation Data for Benchmarking Wire Arc Additive Manufacturing Process Simulations

Residual stresses cause geometric distortion and affect mechanical performance of additively manufactured structures, yet they are notoriously difficult to assess and predict. Distortion (warpage) can drive parts outside dimensional tolerance limits, leading to part rejection or rework. For parts that meet tolerance, locked-in residual stress fields can affect structural integrity during operation, particularly subcritical cracking by fatigue, creep, or corrosion. This work develops benchmark data for a common additive manufacturing process (Wire Arc Additive Manufacturing) that can be applied for calibration and validation of physical process models that predict residual stress fields. The work includes design of two different samples of differing geometry, detailed manufacturing records for a set of physical samples, and an extensive set of residual stress measurement data developed using two diverse techniques (the contour method and neutron diffraction). An initial application of the work is also reported, where a modeling challenge was issued to secure residual stress model predictions from two independent laboratories that were blind to residual stress measurement data. These initial blind residual stress predictions show significant discrepancies relative to the measurement data, illustrating the potential value of the underlying validation data. An open repository for this work, including the sample designs, manufacturing process records, and the residual stress data, is also provided for future application in non-blind validation efforts.

36 MATERIALS SCIENCE↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is continuous generation of an extremely large amount of equipment reliability (ER) data. These data elements come in textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) forms. They provide system engineers with valuable insights and information by discovering anomalous behaviors or degradation trends, identifying possible causes behind such behaviors and trends, and predicting their direct consequences. This paper directly targets the knowledge generation from ER data by putting “data into context.” We employ model-based system engineering (MBSE) of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by first identifying which of the developed MBSE elements they are referring to. This task is harder for textual data since the information contained in issue or maintenance reports needs to be “understood” by a computational tool. We called this process “knowledge extraction” since our methods extract knowledge from textual data. Last, once numeric and textual ER data elements have been processed and “understood,” we discover possible cause-effect relations among them. This is performed by observing whether a logical connection through the MBSE models exists, and if there is a temporal relationship among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 - MATHEMATICS AND COMPUTING↗

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude↗

Fundamental data for modeling electron-induced processes in plasma remediation of perfluoroalkyl substances

Plasma treatment of per- and polyfluoroalkyl substances (PFAS) contaminated water is a potentially energy efficient remediation method. In this treatment, an atmospheric pressure plasma interacts with surface-resident PFAS molecules. Developing a reaction mechanism and modeling of plasma–PFAS interactions requires fundamental data for electron–molecule reactions. In this paper, we present results of electron scattering calculations, potential energy landscapes and their implications for plasma modelling of a dielectric barrier discharge in PFAS contaminated gases, a first step towards modelling of plasma–water–PFAS intereactions. It is found that the plasma degradation of PFAS is dominated by dissociative electron attachment with the importance of other contributing processes varying depending on the molecule. All molecules posses a large number of shape resonances – transient negative ion states – from near-threshold up to ionization threshold. These states lie in the region of the most probable electron energies in the plasma (4–5 eV) and consequently are expected to further enhance the fragmentation dynamics in both dissociative attachment and dissociative excitation.

54 ENVIRONMENTAL SCIENCES↗

RU Net for Automatic Characterization of TRISO Fuel Cross Sections

TRistructural ISOtropic (TRISO) particle fuel is a type of nuclear fuel known for its high-temperature and high-burnup performance. Each sub-millimeter diameter TRISO particle consists of uranium-oxycarbide (UCO) or UO2 fuel kernel, coated with buffer, inner pyrolytic carbon (IPyC), silicon carbide (SiC), and outer pyrolytic carbon (OPyC) layers. The SiC layer acts as the main containment barrier for the TRISO particle to retain the fission products, while the IPyC and OPyC layers provide additional barriers to the release of fission products, especially fission gases. During irradiation, phenomena like kernel swelling, buffer densification, and IPyC fracture may impact fuel performance. Post-irradiation microscopy on entire compact cross sections or samples of individual particles deconsolidated from compacts is often used to identify these irradiation-induced changes in morphology. However, each fuel compact generally contains thousands of TRISO particles. To get statistical information on these phenomena, it is cumbersome work if done manually. For example, to get information about swelling/densification behaviors of different layers or kernels after irradiation, researchers previously manually measured the perimeter of each TRISO layer in hundreds of particles after four rounds of iterative grinding and polishing encompassing more than 2000 cross-section images for a total of four fuel compacts. To attempt to reduce the subjectivity inherent in that process and accelerate data analysis, we conducted a study on the automatic TRISO layer segmentation on cross-sectional microscopic images using Convolutional Neural Networks (CNNs). CNNs are a class of machine learning algorithms specifically designed for processing structured grid data that have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we have generated the large irradiated TRISO layer dataset with more than 2000 cross-section TRISO microscopic images and the corresponding annotated images. Based on these annotated images, we have employed different CNNs for automatic segmentation of different TRISO layers. These include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net has the best performance in terms of intersection-over-union (IoU). Through the aid of these CNN models, we can expedite the analysis of TRISO particle cross-sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

Convolutional Neural Networks↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

On-the-fly data set combinations with RNTuple

With the expected data volume increase for HL-LHC and the even more complex computing challenges set by future colliders, the need for efficient data storage and processing becomes more pressing. ROOT’s next-generation data format and I/O subsystem, RNTTuple, is designed to address these challenges. RNTTuple already demonstrates a clear improvement in storage and I/O efficiency, as well as overall stability and robustness with respect to its predecessor, TTTree. These improvements provide a solid baseline to introduce novel extensions to common high-energy and nuclear physics (HENP) workflows. Notably, many workflows could benefit from the ability to arbitrarily join and chain data set samples at runtime, which could reduce overall storage requirements and improve application runtime and ergonomics. In this paper, we present the RNTupleProcessor, which enables HENP data set combinations with RNTuple. We will discuss the main design considerations, present the interfaces to support data set combinations and show how they integrate in typical workflows.

de Geus, Florine Willemijn [CERN; Twente U., Ensch↗

Informing Robust Functional Relationship Benchmarks: An Evaluation of the Temperature Sensitivity of Ecosystem Respiration Across the Arctic-Boreal Region

During land model development, simulated carbon dynamics are often benchmarked against observational data sets to evaluate model performance. Functional relationship benchmarks are the relationship between a driving variable (e.g., temperature) and a response variable (e.g., ecosystem respiration) and are a promising tool for assessing model performance by evaluating modeled sensitivities to changing environmental conditions. However, observed functional relationships can be influenced by choices made during data collection and throughout the benchmarking process, impacting the inferred skill of land models. To avoid misrepresenting a model's true performance, it is necessary to systematically evaluate best practices when constructing functional relationship benchmarks. We developed a set of guidelines for constructing functional relationship benchmarks, considering the choice of data set, number of daily observations, temporal extent, and temporal resolution across Alaska and Canada over a 20-year period from 2001 to 2020. The temperature sensitivity of ecosystem respiration from observations, evaluated through an apparent Q 10 , is highly variable both spatially and as a result of the data processing approach applied in the benchmark formation. When benchmarking 13 models from the Warming Permafrost Model Intercomparison Project (WrPMIP), the range in inferred model skill is substantially impacted by the choices applied in constructing functional relationship benchmarks. The inferred performance of a given model is most sensitive to the number of daily observations and temporal extent, followed by choice of benchmark data set and temporal averaging. Results from this analysis can guide the development of consistent and robust functional relationships for future model evaluation studies.

Poe, Jeralyn [Northern Arizona University, Flagsta↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is that the amount of equipment reliability (ER) data being continuously generated are extremely large. These data elements come in different forms: textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) and they provide system engineers with valuable insights and information regarding the discovery of anomalous behaviors or degradation trends, the identification of the possible causes behind such behaviors and trends, and the prediction of their direct consequences. This paper directly targets the generation of knowledge from ER data by putting “data into context”. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by identifying first which elements of the developed MBSE elements they are referring to. This task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process “knowledge extraction” where our methods to extract knowledge from textual data. Lastly, once numeric and textual ER data elements have been processed and “understood”, we discover possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if there is a temporal relation among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 MATHEMATICS AND COMPUTING↗