Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Plan Position Indicator Hydrometeor Field Statistics (PPIHYD) Evaluation Data Product Version 1.0

The PPIHYD evaluation data product provides distinct hydrometeor field statistics calculated from U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility scanning radar plan position indicator (PPI) scans. These statistics include the equivalent reflectivity factor and Doppler spectral width percentiles, min/max values, and first four moments (mean, standard deviation, skewness, and kurtosis) of distinct hydrometeor features (clustered hydrometeor fields). Statistics also include morphological properties, water content and precipitation rate parameterization-based estimates, and thermodynamic properties interpolated using the Interpolated Sonde value-added product (INTERPSONDE VAP). The data set is organized in tabular form and is accompanied by mask arrays with corresponding indices. This straightforward file structure simplifies scanning radar data processing and renders this data set useful for process understanding and model evaluation studies. This report describes the data set and its processing algorithm and provides some examples.

54 ENVIRONMENTAL SCIENCES↗

Developing Fluorescence-Based Sensors to Support Rare Earth Element Separation

Rare earth elements (REEs) are essential to most renewable energy technologies. Unfortunately, as we transition to sustainable energy production, the demand for REEs is rapidly growing well beyond current rates of production. As a result, novel means of efficient, scalable, and easily adaptable methods for processing primary and recycle feedstocks are needed. Development and integration of sensors for highly selective in-line monitoring can support more efficient design and testing of such novel separation processes, as well as more cost-effective deployment of those separation flowsheets. Work here will explore the application of fluorescence spectroscopy, a highly sensitive and selective technique, to quantify multiple lanthanides in complex mixtures including known interferents or quenching agents. Results include identification of the optimal excitation wavelength and the limit of detection of various rare earth elements as well as the performance of data-science-based quantification approaches in streams where “unknowns” are present. Overall, the data science tools in conjunction with optical sensor data were able to quantify analytes in the presence of other lanthanides which can be anticipated in the actual industrial stream. Here we include characterization of lanthanides in a microfluidic device similar to those used in new process development. This study demonstrates the capability of utilizing fluorescence spectroscopy to quantify analytes in a complicated solution matrix, suggesting this is a successful approach for in-line monitoring to optimize the separation efficiency in an industrial stream.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FY24 Laboratory Directed Research and Development Annual Report

The Laboratory Directed Research and Development (LDRD) program yields foundational scientific research and development (R&D) essential to growing SRNL’s core competencies, in alignment with SRNL’s Strategic Plan to provide long-term benefits to the Department of Energy (DOE), the National Nuclear Security Administration (NNSA), and other customers and stakeholders. Five strategic goals are outlined in SRNL’s strategic plan: 1) Provide applied science and engineering for EM’s active clean-up sites and LM’s post closure management sites 2) Provide science-based solutions for gaps identified in nonproliferation strategic vision and support the government in activities impacting national security 3) Lead Science, Technology & Engineering as the central technical authority for processing tritium loaded reservoirs and support production of plutonium pits 4) Align science and energy security programs by focusing modern modeling, simulation, and data analytics tools on materials engineering and performance applications 5) Build a workforce for the future

Clark, Sue [Savannah River National Laboratory (SR↗

12th World Conference on Neutron Radiography Program

This program is for the 12th Annual World Neutron Conference hosted by INL, held in Idaho Falls. Included is the conference schedule, abstracts from presenters, outings, a tour of EBR-1, and facilities at MFC including: TREAT, HFEF, and IMCL. INL researcher abstracts are found on the following pgs: 26, 42, 52, 103, 119, 169, 170, 173, and 182. The 12th World Conference on Neutron Radiography (WCNR-12) is an international forum that brings together researchers, engineers, industry practitioners and university students to promote, discuss and disseminate a wide range of topics in neutron imaging. This weeklong conference focuses on the latest methods, instrumentation and facilities, improvements in data processing and interpretation, and new applications of neutron imaging. Ever-improving neutron imaging instruments and new methods push the frontiers of science and address industrial needs. Many user facilities at large research centers with powerful neutron imaging beamlines address the research needs of an increasingly diverse range of applications. Recent developments in accelerator-based sources could expand the user base for neutron imaging by making compact neutron sources available at facilities beyond large neutron sources at major research centers. Industrial practitioners continue to use neutron imaging for practical industrial applications following industry standards. The World Conference on Neutron Radiography series is organized by the International Society for Neutron Radiography (ISNR) to promote the field of neutron imaging and foster interaction and communication between experts and users in the field (see isnr.de). WCNR-12 continues a series of meetings initiated and started in San Diego, California, in December 1981. In intervals of about four years, the conference moved between the United States, Japan and Europe until the most recent one (WCNR-11) in Sydney, Australia, in 2018 (see previous ISNR proceedings here: https://www.isnr.de/index.php/conferences.)

99 - GENERAL AND MISCELLANEOUS↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Resimulation-based self-supervised learning for pretraining physics foundation models

Self-supervised learning (SSL) is at the core of training modern large machine learning models, providing a scheme for learning powerful representations that can be used in a variety of downstream tasks. However, SSL strategies must be adapted to the type of training data and downstream tasks required. We propose resimulation-based self-supervised representation learning (RS3L), a novel simulation-based SSL strategy that employs a method of resimulation to drive data augmentation for contrastive learning in the physical sciences, particularly, in fields that rely on stochastic simulators. By intervening in the middle of the simulation process and rerunning simulation components downstream of the intervention, we generate multiple realizations of an event, thus producing a set of augmentations covering all physics-driven variations available in the simulator. Using experiments from high-energy physics, we explore how this strategy may enable the development of a foundation model; we show how RS3L pretraining enables powerful performance in downstream tasks such as discrimination of a variety of objects and uncertainty mitigation. In addition to our results, we make the RS3L dataset publicly available for further studies on how to improve SSL strategies.

97 MATHEMATICS AND COMPUTING↗

Improving Fission Products at CARIBU (NA-22 Final Report)

Detailed knowledge of fission-product (FP) decay properties is needed for a variety of applications of nuclear science such as nuclear-energy production, nuclear-nonproliferation efforts, nuclear-forensics assessments, and stockpile stewardship, as well as for establishing a comprehensive understanding of the fission process, r-process nucleo-synthesis, and fundamental neutrino science. Although nearly a thousand radioactive isotopes are produced in fission, in many cases key pieces of nuclear data on only a handful of isotopes are needed to make a significant impact.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SITCOMTN-149: An Interim Report on the LSSTComCam On-Sky Campaign

From 24 October to 11 December 2024, the NSF-DOE Vera C. Rubin Observatory conducted an on-sky campaign using the engineering LSST Commissioning Camera (LSSTComCam) to test the end-to-end functionality of hardware and software, as well as operational procedures. This interim report provides a preliminary technical overview of our understanding of the integrated system performance based tests and analyses conducted during the LSSTComCam on-sky campaign. The objectives are to synthesize what we have learned about the system in a timely way to inform on-going commissioning efforts, and to inform the Rubin science community on the progress of the LSSTComCam on-sky campaign. The report is organized into sections to describe major activities during the campaign, as well as multiple aspects of the demonstrated system and science performance. All of the results presented here are to be understood as work in progress using engineering data and the initial versions of the data processing pipelines; the report is a living document that will be updated as analyses are refined.

79 ASTRONOMY AND ASTROPHYSICS↗

OLCF’s Advanced Computing Ecosystem (ACE): FY25 Update for Ongoing Efforts

The advent of widespread use of artificial intelligence (AI) and machine learning (ML) models in science, coupled with fast data production rates of scientific instruments strain the traditional batch-oriented high-performance computing (HPC) environment. As scientific exploration continues to require more data and faster processing and analysis, new emerging technologies and capabilities to enable cross-facility and time-sensitive workflows are required for seamless integration of HPC and experimental facilities. The Advanced Computing Ecosystem (ACE) is a strategic initiative within the Oak Ridge Leadership Computing Facility (OLCF) established in 2024 to support the development of cutting-edge technologies to advance computational research and infrastructure at OLCF and across the Department of Energy (DOE). Several DOE initiatives are spearheading the evolution of the scientific landscape by blurring facility boundaries and connecting the user facilities to advance scientific capabilities and ensure energy dominance. The DOE Integrated Research Infrastructure (IRI) program is one example that is laying a foundation to support complex cross-facility workflows. The IRI program aims to integrate diverse computational resources, data infrastructures, and scientific instruments to facilitate collaboration and accelerate scientific discovery. The Interconnected Science Ecosystem (INTERSECT) initiative at Oak Ridge National Laboratory (ORNL) is another example that aims to revolutionize scientific research through AI-driven, interconnected autonomous laboratories and research facilities. Finally, the American Science Cloud (AmSC), recently announced in the “One Big Beautiful Bill”, aims to leverage prior infrastructure efforts of the IRI and automation and AI efforts of INTERSECT (and others) to build a federated, AI-augmented AmSC platform to unify the DOE’s computing, experimental, and data resources to catalyze scientific innovation.

97 MATHEMATICS AND COMPUTING↗

X-Ray Imaging and Spectroscopy Mission

The X-Ray Imaging and Spectroscopy Mission (XRISM) is a joint mission between the Japan Aerospace Exploration Agency (JAXA) and the National Aeronautics and Space Administration (NASA) in collaboration with the European Space Agency (ESA). In addition to the three space agencies, universities and research institutes from Japan, North America, and Europe have joined to contribute to developing satellite and onboard instruments, data-processing software, and the scientific observation program. XRISM is the successor to the ASTRO-H (Hitomi) mission, which ended prematurely in 2016. Its primary science goal is to examine astrophysical problems with precise, high-resolution X-ray spectroscopy. XRISM promises to discover new horizons in X-ray astronomy. It carries a 6 × 6 pixelized X-ray microcalorimeter on the focal plane of an X-ray mirror assembly (Resolve) and a co-aligned X-ray CCD camera (Xtend) that covers the same energy band over a large field of view. XRISM utilizes the Hitomi heritage, but all designs were reviewed. The attitude and orbit control system was improved in hardware and software. The spacecraft was launched from the JAXA Tanegashima Space Center on 2023 September 6 (UTC). During the in-orbit commissioning phase, the onboard components were activated. Although the gate valve protecting the Resolve sensor with a thin beryllium X-ray entrance window was not yet opened, scientific observation started in 2024 February with the planned performance verification observation program. The nominal observation program commenced with the following guest observation program beginning in 2024 September.

Astronomy and AstroPhysics↗

Characterization of Soil and Rock Magnetic Properties along Multiple Hillslope Transects at Teller Road Site, Seward Peninsula, Alaska, 2018 and 2023

The magnetometer data was collected in multiple directions across the watershed hillslope at the NGEE Arctic Teller Road site at mile marker 27 (TL_MM27) on the Seward Peninsula, Alaska over multiple years in March 2018 and April 2023. The magnetic data were collected using a Geometrics Inc. G-858 gradiometer and G-857 base station in 2018 and the G-864 gradiometer and G857 base station in 2023. The data was collected (in all instances) by towing the gradiometer behind a snow machine around the watershed with the two sensors in a vertical profile with constant spacing during the continuous survey in that specific year. Magnetic total field measurements were collected by gradiometer and base station, and the data processing was performed in Geometrics MagMap2000 software. The processing steps were limited to removal of data spikes (despiking), reading dropouts, and correction/removal of bad GPS points. All offsets between sensors and GPS are stated within the data files and metadata, alongwith the processed and raw data. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).In this data submission there are two sets of raw magnetic data (.bin and .stn for 2018 and base for 2023; raw rover mag for 2023 is in .csv) inside two .zip files that identify the year the mag data was collected. The data are proprietary format to Geometrics and can be opened and processed with MagMap2000 which can be downloaded for free at Geometrics website. There are also two processed data files *.csv for each year and two metadata files *.csv.

54 ENVIRONMENTAL SCIENCES↗

Diaspora: Resilience-enabling services for science from HPC to edge

Scientific applications of interest to DOE must increasingly engage distributed resources (e.g., instruments, remote computers, data stores, edge devices) and deliver more stringent levels of service (e.g., uninterrupted processing of experiment data streams). In such systems, state is distributed and components can fail in many ways, often silently, making application resilience a major concern. Addressing the resilience needs of such applications requires methods for gaining knowledge of resources and applications and for translating that knowledge into action. We are working on addressing these needs in the context of multi-messenger astronomy, where detecting and responding to unusual transient events in multiple cosmic messengers (gravitational wave, electromagnetic, high- energy particles) from different instruments leads to a federated learning problem.

47 OTHER INSTRUMENTATION↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SITCOMTN-161: PSF assessment in the field of Abell 360 and shapeHSM shear profile using LSSTComCam data

The Rubin LSSTComCam on-sky campaign performed at the end of 2024 provided observations of the Abell 360 galaxy cluster; these data allow a preliminary study of cluster weak lensing analysis using Rubin Data Preview 1 (DP1) data. Among all the steps required for such analyses, accurate modeling of the PSF is essential. This work uses several diagnostics, mostly based on the residuals between the second moments of stars and the PSF model, to characterize the accuracy of the PSF modeling in the A360 field. We find the level of the residuals to be sufficiently low not to hinder the measurement of the tangential shear profile around A360. With a simple source selection process, we demonstrate that outputs of the LSST Science Pipelines can be used to detect the tangential shear profile in Abell 360 at the 3.6σ level, and our analysis indicates that contamination from PSF modeling systematics is negligible.

Dell'Antonio, Ian [Brown University]↗

Data Imbalance, Uncertainty Quantification, and Transfer Learning in Data‐Driven Parameterizations: Lessons From the Emulation of Gravity Wave Momentum Transport in WACCM

Abstract Neural networks (NNs) are increasingly used for data‐driven subgrid‐scale parameterizations in weather and climate models. While NNs are powerful tools for learning complex non‐linear relationships from data, there are several challenges in using them for parameterizations. Three of these challenges are (a) data imbalance related to learning rare, often large‐amplitude, samples; (b) uncertainty quantification (UQ) of the predictions to provide an accuracy indicator; and (c) generalization to other climates, for example, those with different radiative forcings. Here, we examine the performance of methods for addressing these challenges using NN‐based emulators of the Whole Atmosphere Community Climate Model (WACCM) physics‐based gravity wave (GW) parameterizations as a test case. WACCM has complex, state‐of‐the‐art parameterizations for orography‐, convection‐, and front‐driven GWs. Convection‐ and orography‐driven GWs have significant data imbalance due to the absence of convection or orography in most grid points. We address data imbalance using resampling and/or weighted loss functions, enabling the successful emulation of parameterizations for all three sources. We demonstrate that three UQ methods (Bayesian NNs, variational auto‐encoders, and dropouts) provide ensemble spreads that correspond to accuracy during testing, offering criteria for identifying when an NN gives inaccurate predictions. Finally, we show that the accuracy of these NNs decreases for a warmer climate (4 × CO 2 ). However, their performance is significantly improved by applying transfer learning, for example, re‐training only one layer using ∼1% new data from the warmer climate. The findings of this study offer insights for developing reliable and generalizable data‐driven parameterizations for various processes, including (but not limited to) GWs.

54 ENVIRONMENTAL SCIENCES↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Development of high throughput and in vitro assays for analyzing RNA modifications

Modifications on RNAs play major roles in their stability, translation, and enzymatic activity. Despite its importance, the current techniques are insufficient to study the structure and function of RNA modifications. Indeed, the National Academies of Science, Engineering and Medicine indicate that developing new tools and further study the function of RNA modifications is strategically a high priority for advancing science in the coming years (https://www.nationalacademies.org/our-work/toward-sequencing-and-mapping-of-rna-modifications). RNA modifications occur in all domains of life controlling processes such as RNA turnover, translation regulation, cellular defenses and bioproduction. Our preliminary data indicated that the insulin mRNA might get ADP-ribosylated by the ADP-ribosyltransferase PARP12. RNA ADP-ribosylation has been described in Escherichia coli. Combined to the fact that ADP-ribosyltransferase (PARP) genes are conserved throughout evolution we hypothesize that this modification might play essential roles in cells. Therefore, we proposed to develop sequencing techniques and in vitro enzymatic assays to identify and validate ADP-ribosylation motifs and sites. Here we report the development of RNA-seq and qPCR assays to identify ADP-ribosylated RNAs, in addition to a nicotinamide adenosine dinucleotide (NAD – ADP-ribosylation donor) consumption assay and an enzyme-linked immunosorbent assay (ELISA) to measure ADP-ribosyltransferase activity. Testing these assays with the insulin mRNA confirmed that this transcript is ADP-ribosylated. These assays will not only enable studying the function of ADP-ribosylation but can be easily adapted for studying other RNA modifications. This will open opportunities to study RNA modifications in different model systems from bacteria to viruses to plants, bringing insights into their cellular functions and the possibility of targeting them for biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES↗