Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

CoCoMET v1.0: a unified open-source toolkit for atmospheric object tracking and analysis

Advances in performance and analysis capabilities have accelerated the development of object tracking algorithms for atmospheric research. This has resulted in a growing number of studies using Lagrangian tracking techniques to analyze the evolution of atmospheric phenomena and the underlying processes. However, the increasing complexity and variety of tracking algorithms present a steep learning curve for new users and make it difficult for existing users to compare algorithm performance. We introduce CoCoMET (Community Cloud Model Evaluation Toolkit), an open-source toolkit that addresses these issues. CoCoMET simplifies the process of running multiple tracking algorithms simultaneously and analyzing objects in both model and observational datasets by specifying parameters in a single configuration file. It standardizes input data from different sources into a consistent format and unifies the tracking output across algorithms. CoCoMET enhances the functionality of existing tracking methods by calculating additional properties such as cell growth and dissipation rates, perimeter, surface area, convexity, and irregularity. In addition, CoCoMET includes a novel method for identifying mergers and splits in 2D and 3D tracks and supports the integration of Eulerian/stationary datasets external to the tracking data for process studies. Its potential utility is demonstrated through examples of model intercomparison, model evaluation against observations, and comparisons between tracking algorithms. Designed for open-source environments, CoCoMET will continue to expand with future releases, incorporating more input data types and tracking algorithms.

54 ENVIRONMENTAL SCIENCES↗

MARVEL Reactor Digital Engineering Developments

The MARVEL reactor project has served to introduce a new generation of engineers to the processes required to transform a reactor design from simply an idea on paper into what will be an approved, constructed, and operational nuclear power system. Much as there have been advances in materials, analysis, and evaluation methodologies over the 50 years since the last reactor was built at INL, so too has the technology for managing the engineering process itself advanced. Digital Engineering tools and methods provide improved coordination between previously siloed engineering disciplines, reduced burdens of non-value-added data transcription processes and bring forward insights and improvements that might otherwise fall later in the design stage, where changes are much more costly. While the tools and techniques to support the full digital engineering vision are not yet complete, the MARVEL design processes provide valuable demonstrations and validations of key aspects and illuminate further areas for implementation by subsequent projects.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Detection of multi-modal Doppler spectra – Part 1: Establishing characteristic signals in radar moment data

Vertically pointing millimeter-wavelength radars provide a wealth of information about cloud and precipitation particle properties. Doppler spectral data can inform on how particles of varying vertical velocities contribute to the total backscattered power observed. It is more computationally cost effective to process moment data instead of spectra data, but doing so leaves valuable information on the cutting room floor. To confidently identify a multi-modal spectra event, in which two or more modes are present within a layer, Doppler spectral data are essential. This means long-term identification of layers featuring multi-modal spectra can be cost prohibitive. To address this, we explore three multi-modal spectra cases from winter precipitation events to determine characteristic signatures of these layers in the moment data averaged over short time periods (∼ 145 s) and explore how these layers differ from the rest of the vertical profiles. We find that the mean spectrum width and the standard deviation of mean Doppler velocity can be used to determine whether or not a layer is multi-modal. In particular, multi-modal layers in mixed-phase and ice clouds feature larger mean spectrum width (exceeding 0.17 m s −1 ) and smaller standard deviation of the mean Doppler velocity (below 0.1 m s −1 ). In Part 1 of this study, the identification criteria and methods are described. In Part 2 (Wugofski and Kumjian, 2025), we perform a verification of the method for three years of vertically pointing radar data, and explore the meteorological conditions associated with identified multi-modal spectral events.

Wugofski, Sarah [Pennsylvania State Univ., Univers↗

Underwater Target Detection Software Demonstration on the RivGen Turbine

This repository contains data and processing scripts necessary to train the object detection models utilized in the underwater target detection software demonstration on the RivGen turbine project and to produce performance metrics (precision, recall, mAP50, mAP50-95). - Contents - Data consist of "images" and "labels". Each image has an associated label, both share the same time string in its file name (e.g., 2024_05_25_09_01_57.98.jpg and 2024_05_25_09_01_57.98.txt). Time strings have the format %yyyy_%mm_%dd_%HH_%MM_%SS.%3f. Images and labels were curated from 2021 and 2024 smolt outmigration periods at the project site in Igiugig, AK. Images are monochrome 8-bit images of objects (smolt, debris, and other) passing through the field of view of the deployed cameras during various operational stages of the RivGen turbine. Labels are text files indicating the class and bounding polygon of each object in an image. The provided labels use the "YOLO" label format. - Requirements - Python3.8+ is required to install and run the train and validation script. The README.md provides instruction for installing the requirements from the requirements.py file. - Instructions - The "example_train.py" file ingests the provided data, trains a model, and produces model performance metrics at completion. NOTE: model performance metrics will vary from run to run as a consequence of the random selection of training and validation data.

16 TIDAL AND WAVE POWER↗

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING↗

Feedstock/Pretreatment Screening for Bioconversion of Sugars and Lignin Residues

This project will conduct biomass deconstruction (pretreatment and enzymatic saccharification) on two representative biomass feedstocks and four different pretreatment processes, including three high temperature steam/chemical pretreatments and one low temperature chemical/mechanical pretreatment. A third biomass feedstock will undergo biomass deconstruction with three different pretreatment processes, including two high temperature chemical pretreatments and one low temperature chemical/mechanical pretreatment. Several pretreatment conditions will be performed in a screening study using NLR pilot-scale pretreatment equipment to generate a range of pretreated biomass substrates. A selected number of these substrates will be chosen for enzymatic saccharification evaluation, based on standard compositional analysis of the pretreated substrates as a primary indicator of pretreatment efficacy. Resulting enzymatic hydrolysis slurries will be analyzed to determine overall biomass sugar yields. Additional compositional analysis will be performed to determine oligomeric sugar composition and structure, to analyze structural characteristics of solids fractions on native, pretreated, and enzymatically saccharified biomass residues for one of the biomass feedstocks, corn stover. The enzymatically saccharified materials will undergo 2,3-butanediol fermentation in a shaker flask as bench scale. Using relevant process performance data collected in these various conversion steps, technoeconomic analysis activities will be performed to compare the economic potential of the various biomass feedstock and pretreatment processes and to identify key economic drivers and sustainability metrics.

09 BIOMASS FUELS↗

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a field grown transgenic poplar engineered with a monolignol 4-O-methyltransferase (MOMT4). The poplar was collected from a 2-year-old rotation trees within a three-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h (in 5 min on and 5 min off cycles to avoid excessive sample heating). The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR experiments. 13C–1H HSQC experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (hsqcetgpsisp2.2) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following acquisition conditions: 220 ppm spectral width in F1 (13C) dimension with 256 data points and 12 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the Bruker’s TopSpin 3.6 software. The NMR spectra provides structural characteristics information about lignin in field grown transgenic MOMT4 poplar. Additional meta data is embedded in the raw spectra figures.

Lignin structure, HSQC, poplar, field trial, MOMT4↗

MAPSTER: Automated Geospatial Data Sharing – Version 1.4.0

The US Department of Energy’s (DOE) Oak Ridge National Laboratory (ORNL) developed MAPSTER which is a geospatial data management tool that aggregates, organizes, and shares data from dispersed sources such as unmanned aerial systems (UAS). Built specifically for use in environments where communications may be limited, MAPSTER utilizes two key technologies to effectively manage data in the field and enable easy data sharing with authorized partners: Observer and Checkpoint. Observer is a lightweight software package on an edge device, such as a laptop, that automatically detects newly processed UAS data and sends to a central server called Checkpoint. Checkpoint is a centralized server at ORNL that receives and manages data from all Observer instances. Even in a very low bandwidth environment, Observer can still send information about the UAS data product almost instantly as it generates its own metadata package on the size of KB (kilobytes). MAPSTER is not only for UAS data but for any geospatial data collected at the austere edge and dispersed sources.

97 MATHEMATICS AND COMPUTING↗

EMPHATIC Silicon Strip Detector Efficiencies

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, Virginia [Illinois U., Urbana (main)]↗

Determining the Efficiency of EMPHATICs Silicon Strip Detectors (SSDs)

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, V. [Illinois U., Urbana (main)]↗

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES↗

White Paper: Scalable Digital Twin Capabilities for Aging and Surveillance of Engineered Systems

This white paper presents a multi-year initiative to develop practical, secure, and scalable digital twin capabilities for engineered systems in aging and surveillance contexts—an approach pioneered at the National Nuclear Security Administration (NNSA) Lawrence Livermore National Laboratory (LLNL) that maps directly onto the needs and ambitions of the Navy for ship- and fleet-level digital twins. LLNL’s work in building part- and process-level digital twins for advanced manufacturing, with a vision to scale up to entire factory floors and, ultimately, enterprise-wide digital twins, offers an adaptable pathway for the Navy as it seeks to modernize lifecycle management, readiness, and predictive maintenance across ships and fleets. For our application, we integrate physics-based modeling with automated data ingestion, processing, and AI-driven calibration, creating hybrid models that are both interpretable and data responsive. We modernized legacy workflows, established centralized data infrastructure, automated experimental pipelines, and demonstrated end-to-end coupling of accelerated aging data with finite element simulations via optimization and surrogate modeling. The result is a generalizable framework that supports part-level digital twins today and lays the groundwork for future system-level twins suitable for Navy applications.

36 MATERIALS SCIENCE↗

Relating Aerial Infrared Thermography Defects to Photovoltaic Performance: Preprint

In this research, we examine the relationship between aerial IR defect analysis and photovoltaic (PV) performance data for twelve utility- and commercial-scale solar sites in the United States. To do this, we fuse the site diagram geoJSON's, aerial infrared thermography (aIRT) defect analyses, and associated inverter time series, allowing for a direct comparison between site defects and time series data. Defect analyses were provided by Zeitview, under its Solar Insights platform. Following the data fusion process, we look at the relationship between system performance and aIRT defects. We investigate the relationship between degradation and hotspot defects, as well as the relationship between AC power data and offline strings and misaligned modules. In general, system degradation was not affected by long-term or balance-of-system (BoS) defects as they occurred infrequently in the data set. However, for one system, a near statistically significant relationship (p-value=0.057) was found when comparing the degradation of inverter blocks with several multi-hotspot defects to all other inverter blocks without this particular defect. There was strong alignment when comparing short-term recoverable module defects such as stuck trackers and offline strings to time series data. In general, we found that when an inverter block has more than 80% of modules flagged for one of these defects, its AC power time data is flat-lined and the inverter block is not producing.

aerial inspection↗

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.

Near Isogeneic Lines↗

Interlaboratory comparison of secondary ion mass spectrometry analysis results from the 7th international technical nuclear forensics working group collaborative material exercise

A review and interlaboratory comparison of secondary ion mass spectrometry (SIMS) data obtained by laboratories participating in the International Technical Nuclear Forensics Working Group (ITWG) 7th Collaborative Material Exercise (CMX-7) has been conducted. The analyzed materials were two uranium compounds in powder form and two pieces of uranium metal. The instruments used in this comparison were a small-geometry (SG) SIMS, CAMECA IMS 7f from the Research Centre Rez, and a large-geometry (LG) SIMS, CAMECA IMS 1280 from Los Alamos National Laboratory. Despite the differences in instruments and analytical procedures, e.g., sample preparation, SIMS setups, and data post-processing, there was good agreement for the 234 U/ 238 U, 235 U/ 238 U, and 236 U/ 238 U ratios in the analyzed materials. The main difference was in the precision, which was, as expected, higher for the LG-SIMS. In addition, a comparison between the laboratories was also made for the image processing algorithms applied to raw data acquired in automated particle measurement (APM) mode. In conclusion, the result of this comparison has led to identification of best practices for setting up parameters of the APM software.

and nuclear chemistry↗

Low latency optical-based mode tracking with machine learning deployed on FPGAs on a tokamak

Active feedback control in magnetic confinement fusion devices is desirable to mitigate plasma instabilities and enable robust operation. Optical high-speed cameras provide a powerful, non-invasive diagnostic and can be suitable for these applications. Here, in this study, we process high-speed camera data, at rates exceeding 100 kfps, on in situ field-programmable gate array (FPGA) hardware to track magnetohydrodynamic (MHD) mode evolution and generate control signals in real time. Our system utilizes a convolutional neural network (CNN) model, which predicts the n = 1 MHD mode amplitude and phase using camera images with better accuracy than other tested non-deep-learning-based methods. By implementing this model directly within the standard FPGA readout hardware of the high-speed camera diagnostic, our mode tracking system achieves a total trigger-to-output latency of 17.6 μs and a throughput of up to 120 kfps. This study at the High Beta Tokamak-Extended Pulse (HBT-EP) experiment demonstrates an FPGA-based high-speed camera data acquisition and processing system, enabling application in real-time machine-learning-based tokamak diagnostic and control as well as potential applications in other scientific domains.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data-Driven Protection Software to classify fault locations by protective zone in distribution systems with high PV penetration

The software contains (a) the source codes to generate Point-on-Wave (PoW) transient data for any feeder model in Alternative Transient Program (ATP) format. Codes provide options to change different steady state settings, including the loading condition and PV capacity and transient state setting like faults type, location and initiation time (b) data post-processing source code to converted data from native format to COMTRADE, csv, HDF5 (c) Docker container to train CNN to classify fault locations by protective zone. The container takes dataset and other training parameters (sampling rate, training epochs, batch size etc) as input to train CNN. The container writes back the trained CNN model, training and testing metrics and plots to the local workstation

Ramesh, Meghana↗

Electromagnetic Induction (EMI) Data, 2024, Trail Creek, Colorado

This dataset contains Electromagnetic Induction (EMI) data collected at Trail Creek, Colorado, in 2024. EMI surveys were conducted to investigate the spatial distribution of electrical conductivity in the subsurface, providing insights into soil moisture and subsurface geological features. The surveys were performed along multiple transects to capture variations in conductivity influenced by changes in soil composition, moisture content, and underlying geological structures. This dataset complements other geophysical data collected in the region, including Electrical Resistivity Tomography (ERT) and Terrestrial LiDAR Scanning (TLS), providing a detailed understanding of the subsurface and its impact on surface vegetation and hydrological processes. The data are valuable for environmental geophysics, ecological research, and hydrological modeling in mountainous ecosystems. The files include: - data.zip: the raw EMI data (.csv) - inversion.zip: the inverted resistivity model (.csv and .kml) - kriging.zip: the kriging resistivity model (.csv, .tif, .kmz) - flmd.csv: file level metadata file describing all files within this dataset - dd.csv: data dictionary file describing the column headers within CSV files This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

CMD Mini-Explorer↗