Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scientific machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

The Zwicky Transient Facility: Data Processing, Products, and Archive

The Zwicky Transient Facility (ZTF) is a new robotic time-domain survey currently in progress using the Palomar 48-inch Schmidt Telescope. ZTF uses a 47 square degree field with a 600 megapixel camera to scan the entire northern visible sky at rates of ∼3760 square degrees/hour to median depths of g ~ 20.8 and r ~ 20.6 mag (AB, 5σ in 30 sec). We describe the Science Data System that is housed at IPAC, Caltech. This comprises the data-processing pipelines, alert production system, data archive, and user interfaces for accessing and analyzing the products. The real-time pipeline employs a novel image-differencing algorithm, optimized for the detection of point-source transient events. These events are vetted for reliability using a machine-learned classifier and combined with contextual information to generate data-rich alert packets. The packets become available for distribution typically within 13 minutes (95th percentile) of observation. Detected events are also linked to generate candidate moving-object tracks using a novel algorithm. Objects that move fast enough to streak in the individual exposures are also extracted and vetted. We present some preliminary results of the calibration performance delivered by the real-time pipeline. The reconstructed astrometric accuracy per science image with respect to Gaia DR1 is typically 45 to 85 milliarcsec. This is the RMS per-axis on the sky for sources extracted with photometric S/N ≥10 and hence corresponds to the typical astrometric uncertainty down to this limit. The derived photometric precision (repeatability) at bright unsaturated fluxes varies between 8 and 25 millimag. The high end of these ranges corresponds to an airmass approaching ∼2—the limit of the public survey. Photometric calibration accuracy with respect to Pan-STARRS1 is generally better than 2%. The products support a broad range of scientific applications: fast and young supernovae; rare flux transients; variable stars; eclipsing binaries; variability from active galactic nuclei; counterparts to gravitational wave sources; a more complete census of Type Ia supernovae; and solar-system objects.

Frank J. Masci↗

On the Need to Align Intent and Implementation in Uncertainty Quantification for Machine Learning

Quantifying uncertainties for machine learning (ML) models is a foundational challenge in modern data analysis. This challenge is compounded by at least two key aspects of the field: (a) inconsistent terminology surrounding uncertainty and estimation across disciplines, and (b) the varying technical requirements for establishing trustworthy uncertainties in diverse problem contexts. In this position paper, we aim to clarify the depth of these challenges by identifying these inconsistencies and articulating how different contexts impose distinct epistemic demands. We examine the current landscape of estimation targets (e.g., prediction, inference, simulation-based inference), uncertainty constructs (e.g., frequentist, Bayesian, fiducial), and the approaches used to map between them. Drawing on the literature, we highlight and explain examples of problematic mappings. To help address these issues, we advocate for standards that promote alignment between the \textit{intent} and \textit{implementation} of uncertainty quantification (UQ) approaches. We discuss several axes of trustworthiness that are necessary (if not sufficient) for reliable UQ in ML models, and show how these axes can inform the design and evaluation of uncertainty-aware ML systems. Our practical recommendations focus on scientific ML, offering illustrative cases and use scenarios, particularly in the context of simulation-based inference (SBI).

Trivedi, Shubhendu [MIT] (ORCID:0000000312374301)↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology↗

Spin-Controllable Dynamics in Defect-Engineered Carbon Nanotubes as Single Photon Emitters: Data-Driven Modeling and Computations

Quantum technologies, such as quantum computing and sensing, require efficient single-photon emission (SPE) sources that operate at room temperature in telecom wavelengths. While several materials can serve as SPE sources, no single platform meets all the criteria for efficiency, ambient operation, and scalability. Single-walled carbon nanotubes (SWCNTs) with covalently attached molecules offer a promising solution. Their SPE can be easily tuned via modifications of the SWCNT's diameter, chirality, and bonded molecules, enabling emission across near-IR to telecom wavelengths at ambient conditions. However, to fully realize the potential of SWCNTs and unlock their quantum capabilities, a deeper understanding of how structural defects from molecular adducts affect their emission and competing photoexcited processes is essential. To address this gap in our knowledge, this project combined quantum chemistry calculations with data-driven methods of cheminformatics (QSAR) and machine learning (ML). The developed computational approaches have provided several design strategies for covalent functionalization of SWCNTs to improve their optical response. The collaboration with Los Alamos National Lab (LANL) enabled direct comparison of computational and experimental data, facilitating method validation. This partnership was enhanced through access to LANL's Center for Integrated Nanotechnologies (CINT) utilizing User Facility Program and summer internships, which provided three NDSU graduate students with hands-on experience at LANL. The outcomes of this project included (1) Advancing the current stage of computational methods in accurate modeling of non-adiabatic spin-dependent photoexcited dynamics and its applicability to nanosystems consisting of thousands of atoms, realized as open-access codes linked to existing DFT-based software; (2) Establishing the relationship between the structure of adducts and SWCNTs and intrinsic excitonic and spin properties of defect states for guiding novel synthetic strategies and experimental probes of chemically functionalized SWCNTs as near-IR emitting materials; (3) Generating virtual libraries of hypothetical functionalized SWCNTs for virtual screening of their chemical structures and optical properties, leveraging new functionalities of SWCNTs; (4) Offering a unique experience for NDSU graduate students that prepared them for future scientific careers related to materials modeling and big data processing. These results were summarized in 12 published journal papers and 3 recently submitted papers. One of a key finding is that the position of defect sites on the SWCNT surface primarily drives the emission redshift (up to 100 meV), while the polarity of the defect-inducing molecules has a much smaller effect (~10 meV). However, the electron-donating or withdrawing properties of a molecule influence selecting reactivity of defect sites. These insights important for optimizing synthetic protocols for desired emissions in SWCNTs. We also revealed that the interaction between two defects at various positions on the SWCNT enhances the redshift and optical activity of states, favoring strong near-IR emission. This suggests that manipulations in defect concentrations is a promising strategy for controlling efficient emission. Mostly important, the defect position was found controllable by the spin states of photoexcited intermediates: Excited aromatic molecules form ortho defects with SWCNTs at their singlet states in the presence of oxygen, while oxygen-free conditions favor para defects via the triplet-state mechanism. Additionally, a heat-activated [2+2] cycloaddition reaction facilitates divalent defect formation with fewer bonding positions that narrows emission bands. These groundbreaking findings have been experimentally validated and significantly advance our understanding of defect chemistry in SWCNTs. Using a novel encoding technique and 3D-MoRSE descriptors, we developed highly accurate ML/QSAR models to predict both the 3D structure and optical properties of SWCNTs with chemical defects. This model enabled the creation of a virtual library of 125,556 structures, providing new insights into the relationship between SWCNT-defect structure and emission.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov↗

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis↗

CLPNets: Coupled Lie–Poisson neural networks for multi-part Hamiltonian systems with symmetries

To accurately compute data-based prediction of Hamiltonian systems, it is essential to utilize methods that preserve the structure of the equations over time. We consider a particularly challenging case of systems with interacting parts that do not reduce to pure momentum evolution. Such systems are essential in scientific computations, such as discretization of a continuum elastic rod, which can be viewed as the group of rotations and translations $SE(3)$. The evolution involves not only the momenta but also the relative positions and orientations of the particles. The presence of Lie group-valued elements, such as relative positions and orientations, poses a problem for applying previously derived methods for data-based computing. We develop a novel method of data-based computation and complete phase space learning of such systems. We follow the original framework of SympNets (Jin et al., 2020) and LPNets (Eldred et al., 2024), building the neural network from phase space mappings that preserve the Lie–Poisson structure. We derive a novel system of mappings that are built into neural networks describing the evolution of such systems. We call such networks Coupled Lie–Poisson Neural Networks, or CLPNets. We consider increasingly complex examples for the applications of CLPNets, starting with the rotation of two rigid bodies about a common axis, progressing to the free rotation of two rigid bodies, and finally to the evolution of two connected and interacting $SE(3)$ components, describing the discretization of an elastic rod into two elements. Our method preserves all Casimir invariants to machine precision, preserves energy to high accuracy, and shows good resistance to the curse of dimensionality, requiring only a few thousand data points for all cases studied (three to eighteen dimensions). Additionally, the method is highly economical in memory requirements, requiring only about 200 parameters for the most complex case considered.

Data-based modeling↗

High Energy Density Physics of Inertial Confinement Fusion Ablator Materials

The historic December 5, 2022 experiment at Lawrence Livermore National Lab’s (LLNL) National Ignition Facility (NIF) reached fusion energy ignition for the first time. This is the most important scientific breakthrough of the 21st century paves the way to future clean inertial fusion energy (IFE). The diamond (high density carbon (HDC)) ablator material used in this experiment displays detrimental effects due to the development of hydrodynamic instabilities at the diamond/fuel interface under shock compression. New alternatives to diamond ablators are required to step up the energy yield in ICF experiments. The unique combination of mechanical strength (approaching that of diamond), the ability to accommodate high-Z dopants (in contrast to diamond), and the tunability of the properties (through synthesis material with varying sp 3 content) make amorphous carbon (a-C) a promising material for next-generation IFE ablative capsules. However, despite its critical importance to the IFE program, the behavior of a-C carbon at extreme temperatures and pressures remains largely unexplored. The primary goals of this project were to perform groundbreaking dynamic compression experiments and predictive simulations to uncover the fundamental high-energy-density physics of amorphous carbon. Our goals were (1) to uncover the metastability range of amorphous carbon and probe phase transitions to diamond or metastable supercooled liquid carbon; (2) to acquire high-quality equation of state (EOS) data and develop an experimentally validated EOS from machine-learning MD simulations of the complex states of carbon; and (3) to uncover the complex behavior of carbon liquid in both thermodynamically stable and metastable supercooled states by accessing large areas of carbon phase diagram with amorphous samples with variable sp 3 content. Our proposed experimental program included measurements of equation of state and diffraction measurements using the Omega EP laser at the Laboratory of Laser Energetics at the University of Rochester. The theoretical/simulation program involved the development of machine-learning models of the complex response of amorphous carbon under dynamic compression by performing molecular dynamics simulations at experimental time and length scales using leadership class DOE supercomputers. Simulations guided experiments to observe predicted phenomena and acquire critical experimental data in specific pressure-temperature domains to validate theoretical models. This research delivered fundamental properties of novel amorphous carbon IFE ablator material, including phase diagram and EOS. These results will aid in IFE target design and implosion experiments. A unique combination of predictive simulations and dynamic and static experiments provided a highly inspirational intellectual environment for graduate students and postdocs involved in this project.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Progressive transfer learning for advancing machine learning-based reduced-order modeling

Abstract To maximize knowledge transfer and improve the data requirement for data-driven machine learning (ML) modeling, a progressive transfer learning for reduced-order modeling (p-ROM) framework is proposed. A key concept of p-ROM is to selectively transfer knowledge from previously trained ML models and effectively develop a new ML model(s) for unseen tasks by optimizing information gates in hidden layers. The p-ROM framework is designed to work with any type of data-driven ROMs. For demonstration purposes, we evaluate the p-ROM with specific Barlow Twins ROMs (p-BT-ROMs) to highlight how progress learning can apply to multiple topological and physical problems with an emphasis on a small training set regime. The proposed p-BT-ROM framework has been tested using multiple examples, including transport, flow, and solid mechanics, to illustrate the importance of progressive knowledge transfer and its impact on model accuracy with reduced training samples. In both similar and different topologies, p-BT-ROM achieves improved model accuracy with much less training data. For instance, p-BT-ROM with four-parent (i.e., pre-trained models) outperforms the no-parent counterpart trained on data nine times larger. The p-ROM framework is poised to significantly enhance the capabilities of ML-based ROM approaches for scientific and engineering applications by mitigating data scarcity through progressively transferring knowledge.

97 MATHEMATICS AND COMPUTING↗

Report for the DOE Office of Science Workshop on Envisioning Frontiers in AI and Computing for Biological Research

Artificial intelligence (AI), machine learning (ML), and high-performance computing (HPC) are poised to transform biological research, spurring innovation in biotechnology and biosystems design. "is transformation will bring an explosion of new capabilities to control the expression of genomic information in living organisms and harness that information to invent new biobased technologies (Jinek et al. 2012; NASEM 2025).

59 BASIC BIOLOGICAL SCIENCES↗

A Science-Focused Artificial Intelligence (AI) Responding in Real-Time to New Information: Capability Demonstration for Ocean World Missions

Introduction: Artificial intelligence (AI) has long been considered a potential mechanism to explore increasingly challenging environments, including those with extreme temperatures and pressures, limited communication capabilities, or those with demanding terrain. We posit that missions in extreme environments could deploy an onboard AI focused on science observations and goals in order to augment a traditional concept(s) of operations (ConOps). An onboard AI capability could perform functions such as data analysis in order to make high-level decisions, including prioritized data transmission for analysis by ground-based teams or autonomously-guided follow-on analyses that maximize science return. Such a capability would empower missions to respond to scientific data of interest in real-time; a mission could make observations and perform a preliminary analysis to alert ground-based scientists to an observation of interest, enabling an informed, rapid response from Earth-based teams. Enceladus Case Study for Onboard AI: We are developing an onboard AI capability for real-time telemetry response that formulates and carries-out informed decisions in service to established mission goals, enabling increased science return of a mission. We focus our AI development for use on a constellation of SmallSats orbiting Enceladus. Our Enceladus case study tests autonomous decision-making capabilities in scenarios with complex orbital dynamics, plume ejecta, extreme cold environments, power restrictions, and a requirement to maximize science return for a potential positive detection of life, while critically evaluating the potential for false positives. Telemetry includes simulated scientific data, spacecraft onboard operational data (e.g., position, velocity, and rotation), and engineering hardware performance data. Enceladus SmallSat Constellation. Our constellation includes eight SmallSat spacecraft in an 8:35 resonant orbit-based formation, leveraging Saturn’s gravitational forces to maintain stable orbits with global coverage around Enceladus. To our knowledge, we simulate the first stable configuration of multiple spacecraft in closed orbits around Enceladus, using a full ephemeris force model (Russell and Lara, 2009). Each spacecraft’s orbit will precess, causing an eastward ground track shift (from an orbiter’s perspective) of each spacecraft for each orbit. However, all spacecraft return to their original positions relative to Enceladus after eight Enceladus revolutions around Saturn. We model communication pathways between SmallSats to understand how information would need to be transmitted across the constellation to enable AI-driven decision-making and resource allocation across the fleet. Capability Demonstration. Our simulated capability demonstration inputs position, velocity, and rotation telemetry from our Enceladus-focused constellation simulations, and mass spectrometry data collected from abiotic and biotic laboratory-analog ocean world experiments (Theiling et al., 2018; Theiling, 2021; Da Poian et al., 2023). Data from these experiments are used to simulate MS measurements and different scenarios of science observations for onboard analysis performed on each of the eight spacecraft. For these demonstrations, we integrate 24 machine learning (ML) algorithms into an onboard intelligence as a ‘knowledge base’, including algorithms evaluating data quality and those predicting (with % confidence) gas composition, ocean aqueous chemistry, and whether the sample was influenced by microbial life. The onboard AI capability is designed to use the knowledge base to come to a consensus-based decision in the interpretation of the observed data in order to request additional action outside of a pre-defined ConOps. Requested actions could include e.g., prioritized downlink to Earth (for analysis by ground-based teams) or follow-on analyses performed across the constellation. The spacecraft’s intelligent onboard planner must then determine whether sufficient resources (e.g., time, power, etc.) are available and weigh the request with mission priorities. In our simulation, the constellation is able to identify potential biosignatures using onboard ML algorithms, evaluate the confidence of that prediction, and perform follow-on analyses across the fleet to confirm the detection, in order to best prepare a transmission of these data to Earth-based teams.

astrobiology↗

Project Development of an Electrochemical Denitration and Caustic Generation System for HLW Pretreatment at Hanford - 26350

An engineering-scale electrochemical processing skid is proposed to perform the denitration of Hanford tank waste, which would help to mitigate a key process concern with the direct feed processing of the Hanford Tank Waste Treatment and Immobilization Plant (WTP). The reduction of nitrates and organic compounds in the waste feed will directly reduce hazardous NOx and ammonia gases generated during the vitrification process, which in turn will aid in addressing potential regulatory and safety challenges associated with processing large volumes of tank waste. This paper highlights the past legacy work, project layout, accomplishments from Phase 1 and research and development envisioned for Phase 2. An innovative electrochemical denitration and caustic generation (EDCGe) process was demonstrated for the pretreatment of tank waste at the Savannah River Site (SRS) in the early 2000s. The denitration electrolyzer, off-gas abatement system, and caustic generator electrolyzer are being developed with the intent that the denitration electrolyzer will convert nitrate and nitrite anions to nitrogen gas while also yielding other gaseous byproducts, which may include N2O, NH3, VOCs, and H2. The gaseous byproducts will be managed via a tandem off-gas catalyst-bed treatment system. The caustic generation electrolyzer will recycle NaOH from the feed to produce a clean caustic stream for use within the batching tanks at Hanford, aiding in the preparation of waste for WTP. The reduction in hazardous emissions and improved waste treatment processes provides a robust solution for nuclear waste management, contributing to environmental safety and regulatory compliance. The EDCGe technology is being adapted, modified, and updated for the preparation of the Direct Feed-High Level Waste (DF-HLW) flowsheet at Hanford. Phase 1 demonstrated a bench-scale proof-of-concept for reactions involving the denitration electrolyzer and gas phase abatement of ammonia. The electrochemical technology is drawing on the scientific outcomes that were reported in the legacy work. The results from Phase 1 demonstrated the viability of the EDCGe system in reducing the nitrogen species of simple non-radioactive waste simulants. Commercially available alloys used as electrode materials and membranes are being studied for the denitration and caustic generation electrolyzers. The continuation of this project holds promise for broader applications, such as energy-efficient ammonia production, and contributes significant advancements in nuclear waste management. Additional material discovery has been investigated into ceramic Na super ion conductive (NaSICON) materials and off-gas abatement catalyst discovery. NaSICON is of interest for selective transport of Na within the electrolyzers to make a clean caustic stream. Future integration and optimization efforts, informed by Phase 1 results and ongoing research, will continue to drive advancements in nuclear waste management technology. The technology developed for the EDCGe treatment of tank waste will also have broader potential to inform other fields, such as energy-efficient ammonia production, as well as ammonia abatement catalysis through the lessons learned in electrochemical nitrate reduction. The applications and benefits of this research extend beyond Hanford and the Savannah River Site, supported by a collaborative team of scientists and engineers from national labs, academia, and industry, ensuring a comprehensive approach to solving complex waste treatment challenges. The team is leveraging advanced electrochemical technologies, machine learning, novel catalysts tailored for gaseous nitrogen species, and cutting-edge reactor systems to enhance the process efficiency and effectiveness of the denitration process.

Rodene, Dylan [Savannah River National Laboratory ↗

Efficient learning of accurate surrogates for simulations of complex systems

Machine learning methods are increasingly deployed to construct surrogate models for complex physical systems at a reduced computational cost. However, the predictive capability of these surrogates degrades in the presence of noisy, sparse or dynamic data. Here, we introduce an online learning method empowered by optimizer-driven sampling that has two advantages over current approaches: it ensures that all local extrema (including endpoints) of the model response surface are included in the training data, and it employs a continuous validation and update process in which surrogates undergo retraining when their performance falls below a validity threshold. We find, using benchmark functions, that optimizer-directed sampling generally outperforms traditional sampling methods in terms of accuracy around local extrema even when the scoring metric is biased towards assessing overall accuracy. Finally, the application to dense nuclear matter demonstrates that highly accurate surrogates for a nuclear equation-of-state model can be reliably autogenerated from expensive calculations using few model evaluations.

79 ASTRONOMY AND ASTROPHYSICS↗

Implementing Scientific Simulation Codes Highly Tailored for Vector Architectures Using Custom Configurable Computing Machines

The motivation for this work comes from an observation that amidst the push for Massively Parallel (MP) solutions to high-end computing problems such as numerical physical simulations, large amounts of legacy code exist that are highly optimized for vector supercomputers. Because re-hosting legacy code often requires a complete re-write of the original code, which can be a very long and expensive effort, this work examines the potential to exploit reconfigurable computing machines in place of a vector supercomputer to implement an essentially unmodified legacy source code. Custom and reconfigurable computing resources could be used to emulate an original application's target platform to the extent required to achieve high performance. To arrive at an architecture that delivers the desired performance subject to limited resources involves solving a multi-variable optimization problem with constraints. Prior research in the area of reconfigurable computing has demonstrated that designing an optimum hardware implementation of a given application under hardware resource constraints is an NP-complete problem. The premise of the approach is that the general issue of applying reconfigurable computing resources to the implementation of an application, maximizing the performance of the computation subject to physical resource constraints, can be made a tractable problem by assuming a computational paradigm, such as vector processing. This research contributes a formulation of the problem and a methodology to design a reconfigurable vector processing implementation of a given application that satisfies a performance metric. A generic, parametric, architectural framework for vector processing implemented in reconfigurable logic is developed as a target for a scheduling/mapping algorithm that maps an input computation to a given instance of the architecture. This algorithm is integrated with an optimization framework to arrive at a specification of the architecture parameters that attempts to minimize execution time, while staying within resource constraints. The flexibility of using a custom reconfigurable implementation is exploited in a unique manner to leverage the lessons learned in vector supercomputer development. The vector processing framework is tailored to the application, with variable parameters that are fixed in traditional vector processing. Benchmark data that demonstrates the functionality and utility of the approach is presented. The benchmark data includes an identified bottleneck in a real case study example vector code, the NASA Langley Terminal Area Simulation System (TASS) application.

Rutishauser, David↗

Human Space Exploration: The Moon, Mars, and Beyond

America is returning to the Moon in preparation for the first human footprint on Mars, guided by the U.S. Vision for Space Exploration. This presentation will discuss NASA's mission, the reasons for returning to the Moon and going to Mars, and how NASA will accomplish that mission in ways that promote leadership in space and economic expansion on the new frontier. The primary goals of the Vision for Space Exploration are to finish the International Space Station, retire the Space Shuttle, and build the new spacecraft needed, to return people to the Moon and go to Mars. The Vision commits NASA and the nation to an agenda of exploration that also includes robotic exploration and technology development, while building on lessons learned over 50 years of hard-won experience. Why the Moon? Many questions about the Moon's potential resources and how its history is linked to that of Earth were spurred by the brief Apollo explorations of the 1960s and 1970s. This new venture will carry more explorers to more diverse landing sites with more capable tools and equipment for extended expeditions. The Moon also will serve as a training ground before embarking on the longer, more difficult trip to Mars. NASA plans to build a lunar outpost at one of the lunar poles, learn to live off the land, and reduce dePendence on Earth for longer missions. America needs to extend its ability to survive in hostile environments close to our home planet before astronauts will reach Mars, a planet very much like Earth. NASA has worked with scientists to define lunar exploration goals and is addressing the opportunities for a range of scientific study on Mars. In order to reach the Moon and Mars within a lifetime and within budget, NASA is building on common hardware, shared knowledge, and unique experience derived from the Apollo Saturn, Space Shuttle and contemporary commercial launch vehicle programs. The journeys to the Moon and Mars will require a variety of vehicles, including the Ares I Crew Launch Vehicle, which transports the Orion Crew Exploration Vehicle, and the Ares V Cargo Launch Vehicle, which transports the Lunar Surface Access Module. The architecture for the lunar missions will use one launch to ferry the crew into orbit, where it will rendezvous with the Lunar Module in the Earth Departure Stage, which will then propel the combination into lunar orbit. The imperative to explore space with the combination of astronauts and robots will be the impetus for inventions such as solar power and water and waste recycling. This next chapter in NASA's history promises to write the next chapter in American history, as well. It will require this nation to provide the talent to develop tools, machines, materials, processes, technologies, and capabilities that can benefit nearly all aspects of life on Earth. Roles and responsibilities are shared between a nationwide Government and industry team. The Exploration Launch Projects Office at the Marshall Space Flight Center manages the design, development, testing, and evaluation of both vehicles and serves as lead systems integrator. A little over a year after it was chartered, the Exploration Launch Projects team is testing engine components, refining vehicle designs, performing wind tunnel tests, and building hardware for the first flight test of Ares I-l, scheduled for spring 2009. The U.S. Vision for Space Exploration lays out a roadmap for a long-term venture of discovery. This endeavor will inspire and attract the best and brightest students to power this nation successfully to the Moon, Mars, and beyond. If one equates the value proposition for space with simple dollars and cents, the potential of the new space economy is tremendous, from orbital space delivery services for the International Space Station to mining and solar energy collection on the Moon and asteroids. The Vision for Space Exploration is fundamentally about bringing the resources of the solar system within the economic sphere of humaind. Given the immense size of our solar system, the amount of available material and energy within it present an enormous economic opportunity.

Sexton, Jeffrey D.↗