Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Machine Learning for Mapping Multipactor Susceptibility in RF Systems: Capabilities and Generalization Constraints

Multipactor is a surface-driven electron avalanche phenomenon that degrades the performance and reliability of radio-frequency (RF) systems in particle accelerator and vacuum electronics applications. Multipactor behavior in a given device structure is conventionally assessed through susceptibility charts, which provide a parameter-space characterization of the instability. In this work, we assess the capabilities of machine-learning (ML) models to learn and predict such susceptibility charts and analyze the constraints governing their generalization across materials. Using a simulation-derived dataset spanning six distinct secondary-electron-yield material profiles in a canonical two-surface planar geometry, we train supervised regression models and artificial neural networks to predict the time-averaged electron growth rate, δavg, across the relevant parameter space. Model performance is evaluated using metrics that explicitly probe the structure of susceptibility charts, including Intersection over Union, Structural Similarity Index, and correlation analysis. Tree-based ensemble models outperform neural-network models in reconstructing susceptibility regions and in generalizing across material domains. Principal-component analysis reveals disjoint material feature distributions, indicating that the piecewise mode structure of multipactor susceptibility is difficult to represent with a single global model and that generalization is constrained by data coverage rather than by model complexity. An exhaustive reduced-coverage study further shows that sparse material-space coverage can yield mean performance in the same general range but producing large variability in the susceptibility-region overlap. These results clarify the capabilities of ML-based surrogate models for parameter-space characterization of multipactor discharge. They also provide guidance for their appropriate use in RF system design.

43 PARTICLE ACCELERATORS↗

Evaluation of Digital Nautical Chart data for confirmation and expansion of GeoNames data

Here, this work examines how Digital Nautical Chart (DNC) data may contribute to the evolution and refinement of GeoNames data for near-shore features. GeoNames features are point data with one or more possible place names. DNC Earth Cover Text (ECRText) objects are map labels positioned nearby their real word counterpart. ECRText feature map position strikes a compromise between association with real features and cartographic readability. This work explores whether ECRText features can confirm (or expand names for) existing locations or contribute new locations through data conflation. Due to name variations and spatial position, conflating these data are nontrivial. Previous work engaged in a brief examination using the trigram string matching algorithm under coarse proximity constraints, indicating that ECRText could provide additional value to GeoNames. This work builds on that study, by engaging in a deeper examination of spatial proximity and exploring conflation agreement across an ensemble of string matching approaches. The result finds strong ensemble agreement about ECRText features which already exist in GeoNames but mixed results about which features contribute new information, as well as exploring why some of these matching techniques fail. With an eye toward automation, computational efficiency was found not to be a constraint in sustaining updates.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

An open retail boundary dataset for South Korea using open data and computer vision technique

Although delineating retail boundaries is important to explore and comprehend the dynamics of the retail sector, it is hard to find studies specifically addressing it in the South Korean context. This study fills this gap by proposing new retail boundaries across South Korea. To achieve this goal, we employed a variety of retailers and building datasets and proposed a unique computer vision-based framework with a deep ensemble voting technique. As a result, we delineated 6,636 distinct retail boundaries that were validated against existing reference retail boundaries. These newly delineated retail boundaries provide valuable insights for researchers, governments, and other relevant stakeholders by enhancing their understanding of retail geography. This dataset can be used as a foundational resource for analyses on topics such as pandemic recovery, retail gentrification, and the resilience of retail spaces in response to e-commerce growth, ultimately contributing to more robust retail sector research in South Korea.

97 MATHEMATICS AND COMPUTING↗

Revealing the Crystalline Architecture of Semicrystalline Ion Exchange Membranes for the Design of Conductive and Durable Alkaline Anion Exchange Membranes

Alkaline anion exchange membrane (AAEM) fuel cells offer a cost-effective alternative to proton exchange membrane (PEM) fuel cells by eliminating the need for expensive precious metal catalysts. In both PEMs and AAEMs, semicrystalline polymers are a common choice, as the crystalline domains can act as mechanical reinforcements that limit swelling and promote mechanical durability in the material. However, spatially resolved characterization of crystalline organization in ion exchange membranes beyond ensemble-averaged X-ray scattering is underrepresented, likely in part due to ionization damage limitations in soft materials. Here, in this study, we resolve the nanometer-size crystallites in semicrystalline ion exchange membranes by applying cryogenic four-dimensional scanning transmission electron microscopy (cryo-4D-STEM) along with data-processing algorithms designed to optimize signals at a low dose to minimize radiation damage. We investigate the effects of synthesis components, including molecular weight and thermal treatment, on a model system of AAEMs in comparison to Nafion, the most commonly used and commercially successful PEM today. We find that excess water uptake in polymer membranes, a property directly associated with weak mechanical durability and with possible negative impacts on ion conductivity, can be reduced by over 30% by varying the polymer's crystalline morphology through changes in synthesis parameters such as molecular weight and thermal history. Our results indicate that this improvement is correlated with smaller crystalline domains with a more homogeneous distribution. More broadly, these results demonstrate how the crystalline architecture of polymer membranes can be tuned through their chemistry and thermal treatment in order to improve their conductivity and durability for commercial fuel cell performance.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING↗

Emulation of the calculations of final r -process abundance patterns with a neural network

This work explores the construction of a fast emulator for the calculation of the final pattern of nucleosynthesis in the rapid neutron capture process (the r-process). An emulator is built using a feed-forward artificial neural network (ANN). We train the ANN with nuclear data and relative abundance patterns. We take as input the β-decay half-lives and the one-neutron separation energy of the nuclei in the rare-earth region. The output is the final isotopic abundance pattern. In this work, we focus on the nuclear data and abundance patterns in the rare-earth region to reduce the dimension of the input and output space. We show that the ANN can capture the effect of the changes in the nuclear physics inputs on the final r-process abundance pattern in the adopted astrophysical conditions. We employ the deep ensemble method to quantify the prediction uncertainty of the neural network emulator. The emulator achieves a speed-up by a factor of about 20 000 in obtaining a final abundance pattern in the rare-earth region. The emulator may be utilized in statistical analyses such as uncertainty quantification, inverse problems, and sensitivity analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

CMIP7 data request: impacts and adaptation priorities and opportunities

The Coupled Model Intercomparison Project Phase 7 (CMIP7) undertook an extensive process to gather community input and refine data requests related to impacts and adaptation applications of Earth System Model (ESM) outputs. The Impacts and Adaptation (I&A) Data Request Team worked with CMIP7 leadership to distribute an open solicitation across many communities that use climate model outputs requesting inputs for new and existing variables, the most applicable temporal characteristics, and groupings of variables that together allow for specific application opportunities. This input was then collated and translated into CMIP7 standard templates for inclusion in the broader data request, leading to 13 I&A data request opportunities, 60 variable groups and 539 unique variables sought by vulnerability, impacts, adaptation, and climate services user communities. Here, we describe these opportunities and variable groups, as well as new insights into how ESM groups can prioritize outputs that set off a chain of further analyses, ultimately informing decisions impacting society and natural systems. These include an emphasis on high-resolution outputs to allow further modeling of climate impacts at regional and local scales, improved representation of extreme weather events, enhanced accuracy of downscaling and bias-adjustment techniques, and support for more detailed assessments for decision-making in adaptation and mitigation strategies. There is also broad interest in more extensive provisioning of two-dimensional variables at the Earth's surface, prioritizing experiments that enhance our understanding of both the recent past and future scenarios, and providing outputs that allow further downscaling and bias adjustment. We emphasize that variable groups are the fundamental level at which to engage with the I&A data request, matching the scale of input and the way output provision enables specific I&A applications. Given resource constraints, we applaud CMIP7 efforts to foster strong engagement and communication between ESM groups and the I&A team to build consensus around prudent compromises in priority variables, temporal resolutions, simulation experiments, time subsets, and ensemble members.

Ruane, Alex C. [NASA Goddard Inst. for Space Studi↗

Predicting critical heat flux with uncertainty quantification and domain generalization using conditional variational autoencoders and deep neural networks

Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Dark Energy Survey Year 6 results: Clustering redshifts and importance sampling of self-organized-maps 𝑛⁡(𝑧) realizations for 3 × 2 ⁢pt samples

This work is part of a series establishing the redshift framework for the 3 × 2 ⁢pt analysis of the Dark Energy Survey Year 6 (DES Y6). For DES Y6, photometric redshift distributions are estimated using self-organizing maps (SOMs), calibrated with spectroscopic and many-band photometric data. To overcome limitations from color-redshift degeneracies and incomplete spectroscopic coverage, we enhance this approach by incorporating clustering-based redshift constraints (clustering-z, or WZ) from angular cross-correlations with BOSS and eBOSS galaxies and eBOSS quasar samples. We define a WZ likelihood and apply importance sampling to a large ensemble of SOM-derived 𝑛⁡(𝑧) realizations, selecting those consistent with the clustering measurements to produce a posterior sample for each lens and source bin. The analysis uses angular scales corresponding to 1.5–5 Mpc to optimize signal-to-noise ratio while mitigating modeling uncertainties and marginalizes over redshift-dependent galaxy bias and other systematics informed by the N-body simulation CARDINAL . While a sparser spectroscopic reference sample limits WZ constraining power at 𝑧 >1.1, particularly for source bins, we demonstrate that combining SOM with WZ improves redshift accuracy and enhances the overall cosmological constraining power of DES Y6. As a result, we estimate an improvement in 𝑆 8 of approximately 10% for cosmic shear and 3 ×2⁢pt analysis, primarily due to the WZ calibration of the source samples.

Cosmological parameters↗

Lattice QCD Study of Pion Electroproduction and Weak Production from a Nucleon

Quantum fluctuations in QCD influence nucleon structure and interactions, with pion production serving as a key probe of chiral dynamics. In this Letter, we present a lattice QCD calculation of multipole amplitudes at threshold, related to both pion electroproduction and weak production from a nucleon, using two gauge ensembles near the physical pion mass. We develop a technique for spin projection and construct multiple operators for analyzing the generalized eigenvalue problem in both the nucleon-pion system in the center-of-mass frame and the nucleon system with nonzero momentum. The numerical lattice results are then compared with those extracted from experimental data and predicted by low-energy theorems incorporating one-loop corrections. Published by the American Physical Society 2025

Gao, Yu-Sheng (ORCID:0009000406829247)↗

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING↗

Dataset for manuscript "Consequences of the failure of equipartition for the p-V behavior of liquid water and the hydration free energy components of a small protein"

Previously, we showed that in the molecular dynamics simulation of a rigid model of water it is necessary to use an integration time-step dt that is less than or equal to 0.5 fs to ensure equipartition between translational and rotational modes. We extended that study in the NVT ensemble to NpT conditions and to an aqueous protein. We study neat liquid water with the rigid, SPC/E model and the protein BBA (PDB ID: 1FME) solvated in the rigid, TIP3P model. We examined integration time-steps ranging from 0.5 fs to 4.0 fs for various thermostat plus barostat combinations. We find that a small time-step, dt, is necessary to ensure consistent prediction of the simulation volume. Hydrogen mass repartitioning alleviates the problem somewhat, but is ineffective for the typical time-step used with this approach. The compressibility, a measure of volume fluctuations, is seen to be sensitive to dt. Using the mean volume estimated from the NpT simulation, we examined the electrostatic and van der Waals contribution to the hydration free energy of the protein in the NVT ensemble. These contributions are also sensitive to dt. In going from a time-step of 2 fs to a time-step of 0.5 fs, the change in the net electrostatic plus van der Waals contribution to the hydration of BBA is already in excess of the folding free energy reported for this protein. The data-set contains the simulation metadata and log files that support the claims noted above.

59 BASIC BIOLOGICAL SCIENCES↗

Generalizable Image Segmentation for Microstructure Characterization Through Integrated SEM and EBSD Analysis

We demonstrate generalizable semantic segmentation using minimal ground truth data. Correlated scanning electron microscopy (SEM) images and electron backscatter diffraction (EBSD) measurements of frictionstir processed 316L stainless steel plates were used to train deep learning models for grain boundary segmentation. Secondary electron (SE) imaging taken at an accelerating voltage of 10 keV correlated to EBSD-derived grain boundaries produced the best performing model. Notably, an ensemble of three models trained on a single SE image produced accurate segmentation over a series of BSE images of samples manufactured under different processing parameters, with a resultant mean absolute error in grain size of 0.34 µm. The striking generalizability of the models likely results from the similar escape depths of the SE training input and the EBSD training output and the reduced probability of dislocation artifacts appearing in the image. This finding highlights the importance of considering the physical principles behind imaging in the development of robust segmentation models for microstructure characterization.

Taufique, Mohammad Fuad Nur↗

GenAI4UQ: A software for forward and inverse uncertainty quantification using conditional generative AI

We introduce GenAI4UQ, a software package for forward and inverse uncertainty quantification in model calibration, parameter estimation, and ensemble forecasting. GenAI4UQ leverages a generative AI-based conditional modeling framework to address limitations of traditional inverse modeling techniques, such as Markov Chain Monte Carlo (MCMC) methods. By replacing computationally intensive iterative processes with a direct, learned mapping, GenAI4UQ enables efficient calibration of input parameters and generation of predictions directly from observations. The software supports rapid ensemble forecasting with robust uncertainty quantification while maintaining computational and storage efficiency. Built-in auto-tuning of hyperparameters simplifies model training, ensuring accessibility for users with varying expertise. Its versatile conditional generative framework is applicable across diverse scientific domains. While GenAI4UQ offers significant advantages in flexibility and efficiency, users should interpret its uncertainty estimates with caution in data-sparse scenarios, as the model may overestimate uncertainty—an effect common to all surrogate-based approaches including MCMC with surrogate models. Despite this, GenAI4UQ transforms inverse modeling by providing a fast, reliable, and user-friendly solution. It empowers researchers and practitioners to quickly estimate parameter distributions and generate model predictions for new observations, facilitating efficient decision-making and advancing the state of uncertainty quantification in computational modeling.

97 MATHEMATICS AND COMPUTING↗

Advanced Polymer Characterization: Modular Operations for Spectral Alignment by Iterative Compression (MOSAIC)

Matrix-assisted laser desorption/ionization (MALDI) mass spectrometry encodes structural information across diverse homo- and copolymer ensembles, yet decrypting these spectra requires a systematic analytical approach. We introduce Modular Operations for Spectral Alignment by Iterative Compression (MOSAIC)─a general cipher algorithm that applies modular arithmetic to filter monomer-derived mass contributions and cluster MALDI peaks by nonconstitutional repeating units (non-CRUs). MOSAIC performs sequential modular operations using monomer mass differences as base units to compress complex spectral data, revealing end-group distributions and comonomer incorporation. As a demonstration, we applied MOSAIC to five copolymers formed by two different polymerization mechanisms. Furthermore, the resulting remainder–mass plots clearly resolve polymer homologs with distinct non-CRUs into visually apparent clusters, enabling intuitive assignment of mass spectral features.

Wang, Hanlin M. [University of Illinois at Urbana−↗

Gas Transfer Across Air‐Water Interfaces in Inland Waters: From Micro‐Eddies to Super‐Statistics

In inland water covering lakes, reservoirs, and ponds, the gas exchange of slightly soluble gases such as carbon dioxide, dimethyl sulfide, methane, or oxygen across a clean and nearly flat air‐water interface is routinely described using a water‐side mean gas transfer velocity $\overline{k_{L}}$, where overline indicates time or ensemble averaging. The micro‐eddy surface renewal model predicts $\overline{k_{L}}$ = α o Sc -1/2 ($v\bar{ϵ}$) 1/4 , where Sc is the molecular Schmidt number, $v$ is the water kinematic viscosity, and $\bar{ϵ}$ is the waterside mean turbulent kinetic energy dissipation rate at or near the interface. While α o = 0.39 - 0.46 has been reported across a number of data sets, others report large scatter or variability around this value range. It is shown here that this scatter can be partly explained by high temporal variability in instantaneous ϵ around $\bar{ϵ}$, a mechanism that was not previously considered. As the coefficient of variation (CV e ) in ϵ increases, α o must be adjusted by a multiplier (1 = CV e 2 ) -3/32 that was derived from a log‐normal model for the probability density function of ϵ. Reported variations in α o with a macro‐scale Reynolds number can also be partly attributed to intermittency effects in ϵ. Such intermittency is characterized by the long‐range (i.e., power‐law decay) spatial auto‐correlation function of ϵ. That α o varies with a macro‐scale Reynolds number does not necessarily violate the micro‐eddy model. Instead, it points to a coordination between the macro‐ and micro‐scales arising from the transfer of energy across scales in the energy cascade.

Batchelor scale↗

Trends and Drivers of Terrestrial Sources and Sinks of Carbon Dioxide: An Overview of the TRENDY Project

The terrestrial biosphere plays a major role in the global carbon cycle, and there is a recognized need for regularly updated estimates of land-atmosphere exchange at regional and global scales. An international ensemble of Dynamic Global Vegetation Models (DGVMs), known as the “Trends and drivers of the regional scale terrestrial sources and sinks of carbon dioxide” (TRENDY) project, quantifies land biophysical exchange processes and biogeochemistry cycles in support of the annual Global Carbon Budget assessments and the REgional Carbon Cycle Assessment and Processes, phase 2 project. DGVMs use a common protocol and set of driving data sets. A set of factorial simulations allows attribution of spatio-temporal changes in land surface processes to three primary global change drivers: changes in atmospheric CO 2 , climate change and variability, and Land Use and Land Cover Changes (LULCC). Here, we describe the TRENDY project, benchmark DGVM performance using remote-sensing and other observational data, and present results for the contemporary period. Simulation results show a large global carbon sink in natural vegetation over 2012–2021, attributed to the CO 2 fertilization effect (3.8 ± 0.8 PgC/yr) and climate (–0.58 ± 0.54 PgC/yr). Forests and semi-arid ecosystems contribute approximately equally to the mean and trend in the natural land sink, and semi-arid ecosystems continue to dominate interannual variability. The natural sink is offset by net emissions from LULCC (–1.6 ± 0.5 PgC/yr), with a net land sink of 1.7 ± 0.6 PgC/yr. Despite the largest gross fluxes being in the tropics, the largest net land-atmosphere exchange is simulated in the extratropical regions.

58 GEOSCIENCES↗

Single-molecule epitranscriptomic analysis of full-length HIV-1 RNAs reveals functional roles of site-specific m6As

Abstract Although the significance of chemical modifications on RNA is acknowledged, the evolutionary benefits and specific roles in human immunodeficiency virus (HIV-1) replication remain elusive. Most studies have provided only population-averaged values of modifications for fragmented RNAs at low resolution and have relied on indirect analyses of phenotypic effects by perturbing host effectors. Here we analysed chemical modifications on HIV-1 RNAs at the full-length, single RNA level and nucleotide resolution using direct RNA sequencing methods. Our data reveal an unexpectedly simple HIV-1 modification landscape, highlighting three predominant N 6 -methyladenosine (m 6 A) modifications near the 3′ end. More densely installed in spliced viral messenger RNAs than in genomic RNAs, these m 6 As play a crucial role in maintaining normal levels of HIV-1 RNA splicing and translation. HIV-1 generates diverse RNA subspecies with distinct m 6 A ensembles, and maintaining multiple of these m 6 As on its RNAs provides additional stability and resilience to HIV-1 replication, suggesting an unexplored viral RNA-level evolutionary strategy.

60 APPLIED LIFE SCIENCES↗