Search NASASearch

SEARCH · Search NASA

Results for “science data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Improving Fission Products at CARIBU (NA-22 Final Report)

Detailed knowledge of fission-product (FP) decay properties is needed for a variety of applications of nuclear science such as nuclear-energy production, nuclear-nonproliferation efforts, nuclear-forensics assessments, and stockpile stewardship, as well as for establishing a comprehensive understanding of the fission process, r-process nucleo-synthesis, and fundamental neutrino science. Although nearly a thousand radioactive isotopes are produced in fission, in many cases key pieces of nuclear data on only a handful of isotopes are needed to make a significant impact.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

SITCOMTN-149: An Interim Report on the LSSTComCam On-Sky Campaign

From 24 October to 11 December 2024, the NSF-DOE Vera C. Rubin Observatory conducted an on-sky campaign using the engineering LSST Commissioning Camera (LSSTComCam) to test the end-to-end functionality of hardware and software, as well as operational procedures. This interim report provides a preliminary technical overview of our understanding of the integrated system performance based tests and analyses conducted during the LSSTComCam on-sky campaign. The objectives are to synthesize what we have learned about the system in a timely way to inform on-going commissioning efforts, and to inform the Rubin science community on the progress of the LSSTComCam on-sky campaign. The report is organized into sections to describe major activities during the campaign, as well as multiple aspects of the demonstrated system and science performance. All of the results presented here are to be understood as work in progress using engineering data and the initial versions of the data processing pipelines; the report is a living document that will be updated as analyses are refined.

79 ASTRONOMY AND ASTROPHYSICS

OLCF’s Advanced Computing Ecosystem (ACE): FY25 Update for Ongoing Efforts

The advent of widespread use of artificial intelligence (AI) and machine learning (ML) models in science, coupled with fast data production rates of scientific instruments strain the traditional batch-oriented high-performance computing (HPC) environment. As scientific exploration continues to require more data and faster processing and analysis, new emerging technologies and capabilities to enable cross-facility and time-sensitive workflows are required for seamless integration of HPC and experimental facilities. The Advanced Computing Ecosystem (ACE) is a strategic initiative within the Oak Ridge Leadership Computing Facility (OLCF) established in 2024 to support the development of cutting-edge technologies to advance computational research and infrastructure at OLCF and across the Department of Energy (DOE). Several DOE initiatives are spearheading the evolution of the scientific landscape by blurring facility boundaries and connecting the user facilities to advance scientific capabilities and ensure energy dominance. The DOE Integrated Research Infrastructure (IRI) program is one example that is laying a foundation to support complex cross-facility workflows. The IRI program aims to integrate diverse computational resources, data infrastructures, and scientific instruments to facilitate collaboration and accelerate scientific discovery. The Interconnected Science Ecosystem (INTERSECT) initiative at Oak Ridge National Laboratory (ORNL) is another example that aims to revolutionize scientific research through AI-driven, interconnected autonomous laboratories and research facilities. Finally, the American Science Cloud (AmSC), recently announced in the “One Big Beautiful Bill”, aims to leverage prior infrastructure efforts of the IRI and automation and AI efforts of INTERSECT (and others) to build a federated, AI-augmented AmSC platform to unify the DOE’s computing, experimental, and data resources to catalyze scientific innovation.

97 MATHEMATICS AND COMPUTING

X-Ray Imaging and Spectroscopy Mission

The X-Ray Imaging and Spectroscopy Mission (XRISM) is a joint mission between the Japan Aerospace Exploration Agency (JAXA) and the National Aeronautics and Space Administration (NASA) in collaboration with the European Space Agency (ESA). In addition to the three space agencies, universities and research institutes from Japan, North America, and Europe have joined to contribute to developing satellite and onboard instruments, data-processing software, and the scientific observation program. XRISM is the successor to the ASTRO-H (Hitomi) mission, which ended prematurely in 2016. Its primary science goal is to examine astrophysical problems with precise, high-resolution X-ray spectroscopy. XRISM promises to discover new horizons in X-ray astronomy. It carries a 6 × 6 pixelized X-ray microcalorimeter on the focal plane of an X-ray mirror assembly (Resolve) and a co-aligned X-ray CCD camera (Xtend) that covers the same energy band over a large field of view. XRISM utilizes the Hitomi heritage, but all designs were reviewed. The attitude and orbit control system was improved in hardware and software. The spacecraft was launched from the JAXA Tanegashima Space Center on 2023 September 6 (UTC). During the in-orbit commissioning phase, the onboard components were activated. Although the gate valve protecting the Resolve sensor with a thin beryllium X-ray entrance window was not yet opened, scientific observation started in 2024 February with the planned performance verification observation program. The nominal observation program commenced with the following guest observation program beginning in 2024 September.

Astronomy and AstroPhysics

Characterization of Soil and Rock Magnetic Properties along Multiple Hillslope Transects at Teller Road Site, Seward Peninsula, Alaska, 2018 and 2023

The magnetometer data was collected in multiple directions across the watershed hillslope at the NGEE Arctic Teller Road site at mile marker 27 (TL_MM27) on the Seward Peninsula, Alaska over multiple years in March 2018 and April 2023. The magnetic data were collected using a Geometrics Inc. G-858 gradiometer and G-857 base station in 2018 and the G-864 gradiometer and G857 base station in 2023. The data was collected (in all instances) by towing the gradiometer behind a snow machine around the watershed with the two sensors in a vertical profile with constant spacing during the continuous survey in that specific year. Magnetic total field measurements were collected by gradiometer and base station, and the data processing was performed in Geometrics MagMap2000 software. The processing steps were limited to removal of data spikes (despiking), reading dropouts, and correction/removal of bad GPS points. All offsets between sensors and GPS are stated within the data files and metadata, alongwith the processed and raw data. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).In this data submission there are two sets of raw magnetic data (.bin and .stn for 2018 and base for 2023; raw rover mag for 2023 is in .csv) inside two .zip files that identify the year the mag data was collected. The data are proprietary format to Geometrics and can be opened and processed with MagMap2000 which can be downloaded for free at Geometrics website. There are also two processed data files *.csv for each year and two metadata files *.csv.

54 ENVIRONMENTAL SCIENCES

Diaspora: Resilience-enabling services for science from HPC to edge

Scientific applications of interest to DOE must increasingly engage distributed resources (e.g., instruments, remote computers, data stores, edge devices) and deliver more stringent levels of service (e.g., uninterrupted processing of experiment data streams). In such systems, state is distributed and components can fail in many ways, often silently, making application resilience a major concern. Addressing the resilience needs of such applications requires methods for gaining knowledge of resources and applications and for translating that knowledge into action. We are working on addressing these needs in the context of multi-messenger astronomy, where detecting and responding to unusual transient events in multiple cosmic messengers (gravitational wave, electromagnetic, high- energy particles) from different instruments leads to a federated learning problem.

47 OTHER INSTRUMENTATION

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

SITCOMTN-161: PSF assessment in the field of Abell 360 and shapeHSM shear profile using LSSTComCam data

The Rubin LSSTComCam on-sky campaign performed at the end of 2024 provided observations of the Abell 360 galaxy cluster; these data allow a preliminary study of cluster weak lensing analysis using Rubin Data Preview 1 (DP1) data. Among all the steps required for such analyses, accurate modeling of the PSF is essential. This work uses several diagnostics, mostly based on the residuals between the second moments of stars and the PSF model, to characterize the accuracy of the PSF modeling in the A360 field. We find the level of the residuals to be sufficiently low not to hinder the measurement of the tangential shear profile around A360. With a simple source selection process, we demonstrate that outputs of the LSST Science Pipelines can be used to detect the tangential shear profile in Abell 360 at the 3.6σ level, and our analysis indicates that contamination from PSF modeling systematics is negligible.

Dell'Antonio, Ian [Brown University]

Data Imbalance, Uncertainty Quantification, and Transfer Learning in Data‐Driven Parameterizations: Lessons From the Emulation of Gravity Wave Momentum Transport in WACCM

Abstract Neural networks (NNs) are increasingly used for data‐driven subgrid‐scale parameterizations in weather and climate models. While NNs are powerful tools for learning complex non‐linear relationships from data, there are several challenges in using them for parameterizations. Three of these challenges are (a) data imbalance related to learning rare, often large‐amplitude, samples; (b) uncertainty quantification (UQ) of the predictions to provide an accuracy indicator; and (c) generalization to other climates, for example, those with different radiative forcings. Here, we examine the performance of methods for addressing these challenges using NN‐based emulators of the Whole Atmosphere Community Climate Model (WACCM) physics‐based gravity wave (GW) parameterizations as a test case. WACCM has complex, state‐of‐the‐art parameterizations for orography‐, convection‐, and front‐driven GWs. Convection‐ and orography‐driven GWs have significant data imbalance due to the absence of convection or orography in most grid points. We address data imbalance using resampling and/or weighted loss functions, enabling the successful emulation of parameterizations for all three sources. We demonstrate that three UQ methods (Bayesian NNs, variational auto‐encoders, and dropouts) provide ensemble spreads that correspond to accuracy during testing, offering criteria for identifying when an NN gives inaccurate predictions. Finally, we show that the accuracy of these NNs decreases for a warmer climate (4 × CO 2 ). However, their performance is significantly improved by applying transfer learning, for example, re‐training only one layer using ∼1% new data from the warmer climate. The findings of this study offer insights for developing reliable and generalizable data‐driven parameterizations for various processes, including (but not limited to) GWs.

54 ENVIRONMENTAL SCIENCES

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES

Development of high throughput and in vitro assays for analyzing RNA modifications

Modifications on RNAs play major roles in their stability, translation, and enzymatic activity. Despite its importance, the current techniques are insufficient to study the structure and function of RNA modifications. Indeed, the National Academies of Science, Engineering and Medicine indicate that developing new tools and further study the function of RNA modifications is strategically a high priority for advancing science in the coming years (https://www.nationalacademies.org/our-work/toward-sequencing-and-mapping-of-rna-modifications). RNA modifications occur in all domains of life controlling processes such as RNA turnover, translation regulation, cellular defenses and bioproduction. Our preliminary data indicated that the insulin mRNA might get ADP-ribosylated by the ADP-ribosyltransferase PARP12. RNA ADP-ribosylation has been described in Escherichia coli. Combined to the fact that ADP-ribosyltransferase (PARP) genes are conserved throughout evolution we hypothesize that this modification might play essential roles in cells. Therefore, we proposed to develop sequencing techniques and in vitro enzymatic assays to identify and validate ADP-ribosylation motifs and sites. Here we report the development of RNA-seq and qPCR assays to identify ADP-ribosylated RNAs, in addition to a nicotinamide adenosine dinucleotide (NAD – ADP-ribosylation donor) consumption assay and an enzyme-linked immunosorbent assay (ELISA) to measure ADP-ribosyltransferase activity. Testing these assays with the insulin mRNA confirmed that this transcript is ADP-ribosylated. These assays will not only enable studying the function of ADP-ribosylation but can be easily adapted for studying other RNA modifications. This will open opportunities to study RNA modifications in different model systems from bacteria to viruses to plants, bringing insights into their cellular functions and the possibility of targeting them for biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES

Bottom-up design of actinide materials from molecular clusters: Demonstration of a general-purpose simulation capability leveraging machine-learned atomic potentials

Actinide thin-film coatings such as uranium dioxide (UO 2 ) play an important role in nuclear reactors and other mission-relevant applications, but realization of their potential requires a deep fundamental understanding of the chemical vapor deposition (CVD) processes used for their growth. The slow experimental progress can be attributed, in part, to the standard safety guidelines associated with handling uranium byproducts, which are often corrosive, toxic, and radioactive. Accurate simulation techniques, when used in concert with experiment, can improve laboratory safety, material durability, and deliverable timeframes. However, state-of-the-art computational methods are either insufficiently accurate or intractably expensive. To remedy this situation, in this project we suggested a machine-learning (ML) accelerated workflow for simulating molecular clustering toward deposition. As a benchmark test case, we considered molecular clustering in steam and assessed independent components of our workflow by comparing with measured thermodynamic properties of water. After analyzing each component individually and finding no fundamental barrier to realization of the workflow, we attempted to integrate the ML component, a Sandia-developed tool called FitSNAP. As this was the first application of FitSNAP to atoms and molecules in the gas phase at Sandia, the method required more fitting data than was originally anticipated. Systematic improvements were made by including in the fit data diatomic potentials, molecular single-bond-breaking curves, and symmetry-constrained intermolecular potentials. We concluded that our strategy provides a feasible pathway toward modeling CVD and related processes, but that extensive training data must be generated before it can be of practical use.

36 MATERIALS SCIENCE

Selenium interaction with iron minerals: Quantitative comparison of sorption and coprecipitation impacts on mobility

Given the significance of selenium (Se) as a micronutrient, the radioactive nature of some of its isotopes, and its affinity to iron (Fe) minerals, extensive research has been conducted on the sorption mechanisms between Se and these minerals. Here, in this study, we employ sorption data sourced from the L-SCIE database and coprecipitation data from available literature to achieve the following objectives: i) establish coherence between adsorption and coprecipitation processes, ii) quantitatively evaluate the importance of these processes in nuclear waste repository science, and iii) propose a forward-looking approach for integrating coprecipitation into reactive transport models. Our findings indicate that a correlation between Se adsorption and coprecipitation can be established using the λ formalism. The comparable log(λ Se(IV) /λ Se(VI) ) ratios derived from adsorption and coprecipitation experiments suggest that these processes can be quantitatively compared and evaluated using our numerical approach. Across all iron oxide phases examined, coprecipitation leads to significantly greater immobilization of Se compared to adsorption. Specifically, for hydrous ferric oxide, hematite, and goethite, coprecipitation is predicted to result in 100–1000 times more Se immobilization compared to adsorption, irrespective of the Se oxidation state (Se(IV) or Se(VI)); notably stronger immobilization potential via coprecipitation was observed for magnetite. The modeling approach and quantitative analysis presented herein clearly highlight the importance of including coprecipitation processes when simulating Se (and other elements) transport, particularly under conditions where mineral compositions are transient or evolving with time. Neglecting coprecipitation in models is likely to lead to significant overestimates of migration.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA

The Journal of Open Source Software (JOSS): Bringing Open-Source Software Practices to the Scholarly Publishing Community for Authors, Reviewers, Editors, and Publishers

Open-source software (OSS) is a critical component of open science, but contributions to the OSS ecosystem are systematically undervalued in the current academic system. The Journal of Open Source Software (JOSS) contributes to addressing this by providing a venue (that is itself free, diamond open access, and all open-source, built in a layered structure using widely available elements/services of the scholarly publishing ecosystem) for publishing OSS, run in the style of OSS itself. A particularly distinctive element of JOSS is that it uses open peer review in a collaborative, iterative format, unlike most publishers. Additionally, all the components of the process—from the reviews to the papers to the software that is the subject of the papers to the software that the journal runs—are open. We describe JOSS’s history and its peer review process using an editorial bot, and we present statistics gathered from JOSS’s public review history on GitHub showing an increasing number of peer reviewed papers each year. We discuss the new JOSSCast and use it as a data source to understand reasons why interviewed authors decided to publish in JOSS. JOSS’s process differs significantly from traditional journals, which has impeded JOSS’s inclusion in indexing services such as Web of Science. In turn, this discourages researchers within certain academic systems, such as Italy’s, which emphasize the importance of Web of Science and/or Scopus indexing for grant applications and promotions. JOSS is a fully diamond open-access journal with a cost of around US$\$$5 per paper for the 401 papers published in 2023. The scalability of running JOSS with volunteers and financing JOSS with grants and donations is discussed.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

A comparative analysis of YOLOv8 and U-Net image segmentation approaches for transmission electron micrographs of polycrystalline thin films

Metallic thin films offer a platform to experimentally study the dynamics of microstructural evolution, but the required transmission electron microscopy (TEM)-based imaging generates complex images that are challenging to segment and quantify. This work provides a comparative analysis of a new YOLOv8 model and an established U-Net model for bright-field TEM images of polycrystals, employing a framework leveraging physical observables to evaluate performance against two hand-traced benchmark datasets. This methodology obviates the comparison of large, diversely structured, and manually labeled datasets that are required to assess performance on a per-image/per-pixel basis. It is found that the YOLOv8 model, adapted for real-time instance segmentation, has up to 43× faster inferencing (NVIDIA GeForce RTX 4090) compared to U-Net and reconstructs hand-traced grain size distributions (GSDs) with excellent fidelity, finding mean diameter within 3% for grains near an optimal magnification; for grains that deviate from the optimal pixel-diameter, the size of small- (large)-diameter grains is systematically over- (under)-estimated. This is partially mitigated by including scale-aware augmentations during training. Moreover, when the bias is corrected post-inference by a rigid shift in distribution, the YOLOv8 model reproduces ground truth GSDs with exceptional fidelity, with statistical tests indicating <5% probability that the distributions are distinct. Based on ground truth data, calibration curves pertaining to this shift can be constructed for a given model. This issue is not present in the U-Net model’s results, indicating that for quantitative measurements where the true size of objects is of interest, special procedures must be implemented for YOLO-based models.

36 MATERIALS SCIENCE

Autonomous hybrid optimization of a SiO 2 plasma etching mechanism

Computational modeling of plasma etching processes at the feature scale relevant to the fabrication of nanometer semiconductor devices is critically dependent on the reaction mechanism representing the physical processes occurring between plasma produced reactant fluxes and the surface, reaction probabilities, yields, rate coefficients, and threshold energies that characterize these processes. The increasing complexity of the structures being fabricated, new materials, and novel gas mixtures increase the complexity of the reaction mechanism used in feature scale models and increase the difficulty in developing the fundamental data required for the mechanism. This challenge is further exacerbated by the fact that acquiring these fundamental data through more complex computational models or experiments is often limited by cost, technical complexity, or inadequate models. In this paper, we discuss a method to automate the selection of fundamental data in a reduced reaction mechanism for feature scale plasma etching of SiO 2 using a fluorocarbon gas mixture by matching predictions of etch profiles to experimental data using a gradient descent (GD)/Nelder–Mead (NM) method hybrid optimization scheme. These methods produce a reaction mechanism that replicates the experimental training data as well as experimental data using related but different etch processes.

36 MATERIALS SCIENCE

Legacy Survey of Space and Time Data Preview 2: deep_coadd dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the deep_coadd dataset type. These are the combination of multiple processed, calibrated, and background- subtracted images, for a patch of sky, for each of the six filters. This release contains 925,460 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS