Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Bottom-up design of actinide materials from molecular clusters: Demonstration of a general-purpose simulation capability leveraging machine-learned atomic potentials

Actinide thin-film coatings such as uranium dioxide (UO 2 ) play an important role in nuclear reactors and other mission-relevant applications, but realization of their potential requires a deep fundamental understanding of the chemical vapor deposition (CVD) processes used for their growth. The slow experimental progress can be attributed, in part, to the standard safety guidelines associated with handling uranium byproducts, which are often corrosive, toxic, and radioactive. Accurate simulation techniques, when used in concert with experiment, can improve laboratory safety, material durability, and deliverable timeframes. However, state-of-the-art computational methods are either insufficiently accurate or intractably expensive. To remedy this situation, in this project we suggested a machine-learning (ML) accelerated workflow for simulating molecular clustering toward deposition. As a benchmark test case, we considered molecular clustering in steam and assessed independent components of our workflow by comparing with measured thermodynamic properties of water. After analyzing each component individually and finding no fundamental barrier to realization of the workflow, we attempted to integrate the ML component, a Sandia-developed tool called FitSNAP. As this was the first application of FitSNAP to atoms and molecules in the gas phase at Sandia, the method required more fitting data than was originally anticipated. Systematic improvements were made by including in the fit data diatomic potentials, molecular single-bond-breaking curves, and symmetry-constrained intermolecular potentials. We concluded that our strategy provides a feasible pathway toward modeling CVD and related processes, but that extensive training data must be generated before it can be of practical use.

36 MATERIALS SCIENCE↗

Selenium interaction with iron minerals: Quantitative comparison of sorption and coprecipitation impacts on mobility

Given the significance of selenium (Se) as a micronutrient, the radioactive nature of some of its isotopes, and its affinity to iron (Fe) minerals, extensive research has been conducted on the sorption mechanisms between Se and these minerals. Here, in this study, we employ sorption data sourced from the L-SCIE database and coprecipitation data from available literature to achieve the following objectives: i) establish coherence between adsorption and coprecipitation processes, ii) quantitatively evaluate the importance of these processes in nuclear waste repository science, and iii) propose a forward-looking approach for integrating coprecipitation into reactive transport models. Our findings indicate that a correlation between Se adsorption and coprecipitation can be established using the λ formalism. The comparable log(λ Se(IV) /λ Se(VI) ) ratios derived from adsorption and coprecipitation experiments suggest that these processes can be quantitatively compared and evaluated using our numerical approach. Across all iron oxide phases examined, coprecipitation leads to significantly greater immobilization of Se compared to adsorption. Specifically, for hydrous ferric oxide, hematite, and goethite, coprecipitation is predicted to result in 100–1000 times more Se immobilization compared to adsorption, irrespective of the Se oxidation state (Se(IV) or Se(VI)); notably stronger immobilization potential via coprecipitation was observed for magnetite. The modeling approach and quantitative analysis presented herein clearly highlight the importance of including coprecipitation processes when simulating Se (and other elements) transport, particularly under conditions where mineral compositions are transient or evolving with time. Neglecting coprecipitation in models is likely to lead to significant overestimates of migration.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Data Analytics for Catalysis Predictions: Are We Ready Yet?

Catalysis informatics has received tremendous attention in recent years as a tool to design catalysts and discover unique descriptors that capture the relationships between chemical properties and catalytic performance. One of the stop-gaps in understanding catalytic effects, which is often ignored and limits the deployment of data science tools, relates to the lack of uniform data. The catalytic cleavage of C–X (X= H, C, N, and O) bonds is relevant to many fundamental catalytic processes. In this Perspective, we performed data analytics on four groups of C–X cleavage reactions that are common in production, upcycling, or reactive separation: the C–C cleavage in cyclopropyl alcohol, the C–H cleavage in hydroacylation reactions, the C–O cleavage in β-O-4 linkages, and the C–N cleavage in amides, using experimental data collected from the literature to understand their underlying correlations. Experimental variables of high impact are identified for each reaction by dimensionality reduction methods. We highlight the urgent need for experimental data sets that include full details on the reaction conditions, such as reagent concentration, reaction temperature, or time in machine-readable forms. We discuss the potential improvement of the data of these reactions and promising approaches such as autonomous experiments to fill the gaps in unbiased experimental data. Finally, we also address the early stage consideration of separation aspects in the experimental design of efficient catalytic systems for these fundamental examples of chemical reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Journal of Open Source Software (JOSS): Bringing Open-Source Software Practices to the Scholarly Publishing Community for Authors, Reviewers, Editors, and Publishers

Open-source software (OSS) is a critical component of open science, but contributions to the OSS ecosystem are systematically undervalued in the current academic system. The Journal of Open Source Software (JOSS) contributes to addressing this by providing a venue (that is itself free, diamond open access, and all open-source, built in a layered structure using widely available elements/services of the scholarly publishing ecosystem) for publishing OSS, run in the style of OSS itself. A particularly distinctive element of JOSS is that it uses open peer review in a collaborative, iterative format, unlike most publishers. Additionally, all the components of the process—from the reviews to the papers to the software that is the subject of the papers to the software that the journal runs—are open. We describe JOSS’s history and its peer review process using an editorial bot, and we present statistics gathered from JOSS’s public review history on GitHub showing an increasing number of peer reviewed papers each year. We discuss the new JOSSCast and use it as a data source to understand reasons why interviewed authors decided to publish in JOSS. JOSS’s process differs significantly from traditional journals, which has impeded JOSS’s inclusion in indexing services such as Web of Science. In turn, this discourages researchers within certain academic systems, such as Italy’s, which emphasize the importance of Web of Science and/or Scopus indexing for grant applications and promotions. JOSS is a fully diamond open-access journal with a cost of around US$\$$5 per paper for the 401 papers published in 2023. The scalability of running JOSS with volunteers and financing JOSS with grants and donations is discussed.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A comparative analysis of YOLOv8 and U-Net image segmentation approaches for transmission electron micrographs of polycrystalline thin films

Metallic thin films offer a platform to experimentally study the dynamics of microstructural evolution, but the required transmission electron microscopy (TEM)-based imaging generates complex images that are challenging to segment and quantify. This work provides a comparative analysis of a new YOLOv8 model and an established U-Net model for bright-field TEM images of polycrystals, employing a framework leveraging physical observables to evaluate performance against two hand-traced benchmark datasets. This methodology obviates the comparison of large, diversely structured, and manually labeled datasets that are required to assess performance on a per-image/per-pixel basis. It is found that the YOLOv8 model, adapted for real-time instance segmentation, has up to 43× faster inferencing (NVIDIA GeForce RTX 4090) compared to U-Net and reconstructs hand-traced grain size distributions (GSDs) with excellent fidelity, finding mean diameter within 3% for grains near an optimal magnification; for grains that deviate from the optimal pixel-diameter, the size of small- (large)-diameter grains is systematically over- (under)-estimated. This is partially mitigated by including scale-aware augmentations during training. Moreover, when the bias is corrected post-inference by a rigid shift in distribution, the YOLOv8 model reproduces ground truth GSDs with exceptional fidelity, with statistical tests indicating <5% probability that the distributions are distinct. Based on ground truth data, calibration curves pertaining to this shift can be constructed for a given model. This issue is not present in the U-Net model’s results, indicating that for quantitative measurements where the true size of objects is of interest, special procedures must be implemented for YOLO-based models.

36 MATERIALS SCIENCE↗

Autonomous hybrid optimization of a SiO 2 plasma etching mechanism

Computational modeling of plasma etching processes at the feature scale relevant to the fabrication of nanometer semiconductor devices is critically dependent on the reaction mechanism representing the physical processes occurring between plasma produced reactant fluxes and the surface, reaction probabilities, yields, rate coefficients, and threshold energies that characterize these processes. The increasing complexity of the structures being fabricated, new materials, and novel gas mixtures increase the complexity of the reaction mechanism used in feature scale models and increase the difficulty in developing the fundamental data required for the mechanism. This challenge is further exacerbated by the fact that acquiring these fundamental data through more complex computational models or experiments is often limited by cost, technical complexity, or inadequate models. In this paper, we discuss a method to automate the selection of fundamental data in a reduced reaction mechanism for feature scale plasma etching of SiO 2 using a fluorocarbon gas mixture by matching predictions of etch profiles to experimental data using a gradient descent (GD)/Nelder–Mead (NM) method hybrid optimization scheme. These methods produce a reaction mechanism that replicates the experimental training data as well as experimental data using related but different etch processes.

36 MATERIALS SCIENCE↗

Legacy Survey of Space and Time Data Preview 2: deep_coadd dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the deep_coadd dataset type. These are the combination of multiple processed, calibrated, and background- subtracted images, for a patch of sky, for each of the six filters. This release contains 925,460 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 2: source dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the source dataset type. These are measurements for detected sources in processed visit images. This release contains 28,589 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder↗

Stream discharge and temperature data collected within the East and Taylor Watershed, Colorado for the Lawrence Berkeley National Laboratory Watershed Function Science Focus Area (water years 2019 to 2025)

This dataset contains stream discharge and temperature data for water years 2019 to 2025 from the East and Taylor Watersheds in Colorado, United States. This data was collected to understand hydrological processes occurring in the East River and Taylor River Watersheds, Colorado, which is part of the Lawrence Berkeley National Laboratory Watershed Function Scientific Focus Area. Data includes instantaneous observed discharge using salt dilution and acoustic doppler velocimeter techniques, raw pressure transducer downloaded data, sub-hourly temperature as well as corrected water level and associated stream discharge and mean daily values. Notes on water level corrections, rating curve development and metadata provided. A rating curve is the translation of depth to streamflow. The rating curve can be used as a quantitative measure of the “quality of the data.” Data within this dataset is formatted using ESS-DIVE’s Hydrological Monitoring Reporting Format. This data package contains (1) a zip file (Stream_Discharge_Data_WY19-WY25.zip) containing stream discharge and temperature data organized by location; (2) an InstallationMethods file (InstallationMethods.csv) describing metadata about the installation; (3) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; (4) a data dictionary (dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (5) a locations metadata file (locations.csv); (6) and a sensor metadata file (sensors.csv). All data files are in non-proprietary formats (csv, png, or pdf formats). Please contact Rosemary Carroll, Curtis Beutler, or Austin Shirley for any support in accessing the files. Update on 2023-05-12: Additional data from WYs 2021 and 2022 were added. Additionally, the dataset was converted using ESS-DIVE’s Hydrological Monitoring Reporting Format. Data files were reformatted to match reporting format guidance, new metadata files were added, and files were converted from excel to CSV. Update on 2025-05-16: Additional data from WYs 2022 (for locations not previously included), 2023, and 2024 were added. An additional descriptive PDF (WFSFA_Streamflow_Hydrograph_Disclaimer.pdf) was added. Metadata files were updated to reflect the addition of new data and locations. Update on 2026-05-18: Additional data from WY 2025 were added, including a new location Upper Trail Creek (TR-TCG2). Metadata files were updated to reflect the addition of new data.

54 ENVIRONMENTAL SCIENCES↗

Feature Based Qualification (FBQ) of Wire Arc Additively Manufactured (WAAM) 17-4PH Martensitic Stainless Steels

The Department of Defense (DOD) programs of records desire to reduce the time and cost of the development and delivery loop in metal additive manufacturing (AM), including establishing forwarded AM capabilities. The success of these efforts relies on a robust and qualified process. To achieve this, the United States Army Combat Capabilities Development Command Ground Vehicle Systems Center Materials Engineering (GVME) needs to be able to quickly evaluate, test, and develop feedstocks, processes, and parts. This report is directed towards demonstrating the need for a framework for metal AM processes, defining and exploring geometries for metal AM process qualification, testing resultant deposition, and delivering actionable data.

36 MATERIALS SCIENCE↗

Human Host Cellular Response to HCoV-229E Infection Transcriptomics (ACS-DP1)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5), immortalized human lung fibroblasts cells (MRC5) (MOI5), and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue. Sample data was acquired using an Illumina HiSeq 2000 sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Legacy Survey of Space and Time Data Preview 2: Source searchable catalog

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of a searchable catalog named Source. This catalog contains measurements for detected sources in processed visit images. This catalog contains 17,566,180,086 rows with 154 columns.

79 ASTRONOMY AND ASTROPHYSICS↗

CROCUS Weather Data at Argonne National Laboratory Prairie Site

Vaisala WXT sensor is an all-in-one weather instrument that provides 6 of the most important weather parameters: barometric pressure, temperature, relative humidity, rainfall, wind speed and direction. Temperature, pressure, relative humidity, and rainfall are sampled at 1 second frequency, while wind speed/direction is measured at ten per second (10Hz) frequency. These measurements are useful for looking at characterizing local weather, identifying unique weather events, and studying local turbulence, especially given the high temporal resolution of the wind measurements. These measurements are collected at the Argonne Testbed for Multiscale Observational Science (ATMOS), a prairie field site at Argonne National Laboratory in Lemont, Illinois. Data is available in the netCDF data format, we encourage data users review documentation through Project Pythia to understand how to work with netCDF data https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html. The data is aggregated into daily frequency to make it easier to process multiple days, and compress the higher-resolution fields. Each file contains one day's worth of data (24 hours, starting at 0000 UTC). File naming convention includes the project (CROCUS), location (atmos), data level (raw, a1), date (year, month, day), and hour (0000).

54 ENVIRONMENTAL SCIENCES↗

BLOC Site - ASSIST Thermodynamic Retrievals TROPoe v0.18 / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). This is a post-processed dataset and recommended for use. The profiles are retrieved every 10 minutes from instantaneous radiances observed with an Atmospheric Sounder Spectrometer by Infrared Spectral Technology (ASSIST, Michaud-Belleau et al. 2025) operated by NOAA Physical Sciences Laboratory (PSL) on Block Island for WFIP3. The spectral bands used in the retrieval are in the wavenumber range from 612 - 905.4 cm-1 and are specified in Turner and Löhnert (2021). Additional input data in TROPoe are cloud base height from a collocated ceilometer operated by NOAA GML and temperature, water vapor mixing ratio, and pressure from a collocated surface tower operated by NOAA PSL. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Upton, NY. The TROPoe docker container (version 0.18) is available from Docker Hub at https://hub.docker.com/r/davidturner53/tropoe/tags, and the source code code is available in the GitHub repository https://github.com/OAR-atmospheric-observations/TROPoe.

17 WIND ENERGY↗

NANT Site - ASSIST Thermodynamic Retrievals TROPoe v0.18 / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). This is a post-processed dataset and recommended for use. The profiles are retrieved every 10 minutes from instantaneous radiances observed with an Atmospheric Sounder Spectrometer by Infrared Spectral Technology (ASSIST, Michaud-Belleau et al. 2025) operated by NOAA Physical Sciences Laboratory (PSL) on Nantucket Island for WFIP3. The spectral bands used in the retrieval are in the wavenumber range from 612 - 905.4 cm-1 and are specified in Turner and Löhnert (2021). Additional input data in TROPoe are cloud base height from a collocated ceilometer operated by NOAA GML and temperature, water vapor mixing ratio, and pressure from a collocated surface tower operated by NOAA PSL. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Upton, NY. The TROPoe docker container (version 0.18) is available from Docker Hub at https://hub.docker.com/r/davidturner53/tropoe/tags, and the source code code is available in the GitHub repository https://github.com/OAR-atmospheric-observations/TROPoe.

17 WIND ENERGY↗

High-count rate effects in event processing for XRISM/ Resolve X-ray microcalorimeter: I. Ground test

The spectroscopic performance of an X-ray microcalorimeter is compromised at high count rates. We utilize the Resolve X-ray microcalorimeter onboard the XRISM satellite to examine the effects observed during high-count rate measurements and propose modeling approaches to mitigate them. We specifically address the following instrumental effects that impact performance: CPU limit, pile-up, and untriggered electrical cross-talk. Experimental data at high count rates were acquired during ground testing using the flight model instrument and a calibration X-ray source. In the experiment, data processing not limited by the performance of the onboard CPU was run in parallel, which cannot be done in orbit. This makes it possible to access the data degradation caused by limited CPU performance. We use these data to develop models that allow for a more accurate estimation of the aforementioned effects. To illustrate the application of these models in observation planning, we present a simulated observation of GX 13+1. Understanding and addressing these issues is crucial to enhancing the reliability and precision of X-ray spectroscopy in situations characterized by elevated count rates.

47 OTHER INSTRUMENTATION↗