Search NASA⌕ Search

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Systematic engineering for production of anti-aging sunscreen compound in Pseudomonas putida

Sunscreen has been used for thousands of years to protect skin from ultraviolet radiation. However, the use of modern commercial sunscreen containing oxybenzone, ZnO, and TiO 2 has raised concerns due to their negative effects on human health and the environment. In this study, we aim to establish an efficient microbial platform for production of shinorine, a UV light absorbing compound with anti-aging properties. First, we methodically selected an appropriate host for shinorine production by analyzing central carbon flux distribution data from prior studies alongside predictions from genome-scale metabolic models (GEMs). We enhanced shinorine productivity through CRISPRi-mediated downregulation and utilized shotgun proteomics to pinpoint potential competing pathways. Simultaneously, we improved the shinorine biosynthetic pathway by refining its design, optimizing promoter usage, and altering the strength of ribosome binding sites. Finally, we conducted amino acid feeding experiments under various conditions to identify the key limiting factors in shinorine production. The study combines meta-analysis of 13 C-metabolic flux analysis, GEMs, synthetic biology, CRISPRi-mediated gene downregulation, and omics analysis to improve shinorine production, demonstrating the potential of Pseudomonas putida KT2440 as platform for shinorine production.

59 BASIC BIOLOGICAL SCIENCES↗

GPU-Accelerated Solution of the Bethe–Salpeter Equation for Large and Heterogeneous Systems

We present a massively parallel GPU-accelerated implementation of the Bethe–Salpeter equation (BSE) for the calculation of the vertical excitation energies (VEEs) and optical absorption spectra of condensed and molecular systems, starting from single-particle eigenvalues and eigenvectors obtained with density functional theory. The algorithms adopted here circumvent the slowly converging sums over empty and occupied states and the inversion of large dielectric matrices through a density matrix perturbation theory approach and a low-rank decomposition of the screened Coulomb interaction, respectively. Further computational savings are achieved by exploiting the nearsightedness of the density matrix of semiconductors and insulators to reduce the number of screened Coulomb integrals. We scale our calculations to thousands of GPUs with a hierarchical loop and data distribution strategy. The efficacy of our method is demonstrated by computing the VEEs of several spin defects in wide-band-gap materials, showing that supercells with up to 1000 atoms are necessary to obtain converged results. We discuss the validity of the common approximation that solves the BSE with truncated sums over empty and occupied states. In conclusion, we then apply our GW-BSE implementation to a diamond lattice with 1727 atoms to study the symmetry breaking of triplet states caused by the interaction of a point defect with an extended line defect.

Absorption spectra↗

Regularization via f -Divergence: An Application to Multi-Oxide Spectroscopic Analysis

In this paper, we explore the application of convolutional neural networks (CNNs) for predicting the chemical composition of complex geologic samples in a simulated Martian atmospheric environment. Specifically, we aim to characterize oxide weight percentages (wt.%) of rock samples analyzed by remote Laser-Induced Breakdown Spectroscopy (LIBS), framing the problem as a multi-target regression task . Neural networks trained on LIBS spectra are prone to overfitting due to high spectral complexity, limited labeled data, and measurement noise. While regularization is critical for improving generalization, common methods (e.g., ℓ 2 regularization) impose constraints not directly tied to data distribution properties. We propose a novel regularization method based on a specific ƒ-divergence induced by a graph-based estimator, designed to constrain the distributional discrepancy between predictions and targets. This regularizer serves a dual purpose: (a) mitigating overfitting by enforcing a constraint on the distributional difference between predictions and noisy targets, and (b) acting as an auxiliary loss that penalizes large divergences. To enable backpropagation, we develop a differentiable approximation of this particular ƒ-divergence, making the method feasible for neural networks. Experiments on ChemCam and SuperCam LIBS calibration spectra show that mathematical equation-divergence regularization outperforms or matches standard regularization methods (ℓ 1 , ℓ 2 , dropout) and the classical baseline, partial least squares (PLS). Combining ƒ-divergence regularization with standard regularization yields further performance gains, indicating that distributional regularization is useful in this context giving a promising direction for robust model training in planetary science applications. Source code is publicly available at Klein and Li (2025), https://doi.org/10.11578/dc.20250530.7.

58 GEOSCIENCES↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Measurement of the small-scale 3D Lyman- α forest power spectrum

Small-scale correlations measured in the Lyman-α (Lyα) forest encode information about the intergalactic medium and the primordial matter power spectrum. In this article, we present and implement a simple method to measure the 3-dimensional power spectrum, P 3D , of the Lyα forest at wavenumbers k corresponding to small, ~ Mpc scales. In order to estimate P 3D from sparsely and unevenly distributed data samples, we rely on averaging 1-dimensional Fourier Transforms, as previously carried out to estimate the 1-dimensional power spectrum of the Lyα forest, P 1D . Further, this methodology exhibits a very low computational cost. We confirm the validity of this approach through its application to Nyx cosmological hydrodynamical simulations. Subsequently, we apply our method to the eBOSS DR16 Lyα forest sample, providing as a proof of principle, a first P 3D measurement averaged over two redshift bins z = 2.2 and z = 2.4. This work highlights the potential for forthcoming P 3D measurements, from upcoming large spectroscopic surveys, to untangle degeneracies in the cosmological interpretation of P 1D .

79 ASTRONOMY AND ASTROPHYSICS↗

Accurate field-level weak lensing inference for precision cosmology

We present miko, a catalog-to-cosmology pipeline for general flat-sky field-level inference, which provides access to cosmological information beyond the two-point statistics. In the context of weak lensing, we identify several new field-level analysis systematics (such as aliasing, Fourier mode-coupling, and density-induced shape noise), quantify their impact on cosmological constraints, and correct the biases to a percent level. Next, we find that model misspecification can lead to both absolute bias and incorrect uncertainty quantification for the inferred cosmological parameters in realistic simulations. The Gaussian map prior infers unbiased cosmological parameters, regardless of the true data distribution, but it yields overconfident uncertainties. The log-normal map prior quantifies the uncertainties accurately, although it requires careful calibration of the shift parameters for unbiased cosmological parameters. Here, we demonstrate systematics control down to the 2% level for both models, making them suitable for ongoing weak lensing surveys.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Contrastive learning for robust representations of neutrino data

In neutrino physics, analyses often depend on large simulated datasets, making it essential for models to generalize effectively to real-world detector data. Contrastive learning, a well-established technique in deep learning, offers a promising solution to this challenge. By applying controlled data augmentations to simulated data, contrastive learning enables the extraction of robust and transferable features. This improves the ability of models trained on simulations to adapt to real experimental data distributions. In this paper, we investigate the application of contrastive learning methods in the context of neutrino physics. Through a combination of empirical evaluations and theoretical insights, we demonstrate how contrastive learning enhances model performance and adaptability. Additionally, we compare it to other domain adaptation techniques, highlighting the unique advantages of contrastive learning for this field. Published by the American Physical Society 2025

Wilkinson, Alex (ORCID:0000000253404506)↗

FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy Compression

Cross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. Here, to bridge this gap, we propose FedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show that FedEFsz improves the fairness across different benchmarks by up to 60.88% and meanwhile reduces the communication traffic by up to 315×.

Cross-Silo Federated Learning Systems↗

DP-TwoLevel: two-stage gradient subspace learning for differentially private federated learning

Federated learning (FL) enables collaborative model training across distributed data sources without sharing raw data, but faces fundamental challenges in communication efficiency and privacy. Differentially private (DP) training mitigates information leakage but introduces noise that degrades model performance, especially in high-dimensional settings. We propose DP-TwoLevel, a hierarchical gradient projection method that improves utility under fixed DP constraints by exploiting low-dimensional structure in model updates. Our approach learns a two-level PCA-based representation of gradients and applies DP noise in a reduced-dimensional subspace, thereby lowering the effective noise magnitude while preserving dominant signal components. We evaluate the method across three datasets (MNIST, Fashion-MNIST, CIFAR-10) and three privacy regimes (ϵ∈0.5, 1.0, 2.0). Across nine experimental settings, DP-TwoLevel consistently outperforms DP-FedAvg, achieving an average accuracy improvement of 9.44%, with larger gains observed in lower ϵ(higher-noise) regimes (up to +22.31%). We further analyze scalability across models ranging from 100K to 1.49M parameters and identify a variance-based success criterion: performance remains strong when the projection preserves more than 75% of gradient variance, degrades in a marginal regime (65–75%), and fails below this threshold. Our results demonstrate that structure-aware dimensionality reduction can significantly improve the privacy–utility tradeoff in FL without modifying formal privacy guarantees. We also provide empirical evidence of scaling limitations for global projections and motivate per-layer extensions for larger models.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DDCP framework

DDCP protocol software 1.0 This repository contains the C++ implementation of version 1.x of the Distributed Data Communications Protocol (DDCP). DDCP provides request/reply, feature discovery, data transfer, control, interrupt, and transaction support for communicating with accelerator instrumentation over UDP. The standard server port is 65000. The framework is a source dependency for services that communicate directly with DDCP hardware. It is not a deployable service by itself.

Joshi, Shreya [Fermi National Accelerator Laborato↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Iterative HOMER with uncertainties

We present iHOMER, an iterative version of the HOMER method to extract Lund fragmentation functions from experimental data. Through iterations, we address the information gap between latent and observable phase spaces and systematically remove bias. To quantify uncertainties on the inferred weights, we use a combination of Bayesian neural networks and uncertainty-aware regression. We find that the combination of iterations and uncertainty quantification produces well-calibrated weights that accurately reproduce the data distribution. A parametric closure test shows that the iteratively learned fragmentation function is compatible with the true fragmentation function.

Butter, Anja [Heidelberg Univ. (Germany); Sorbonne↗

Optimal Transport as a Tool for Scientific Discovery in Radiation Biology

This report summarizes findings from research conducted for the “Exploration of the Poten tial for Artificial Intelligence and Machine Learning to Advance Low-Dose Radiation Biology Re search” (RadBio-AI) program, supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, under Awards KP1601011/FWP CC121 and KP1601017/FWP CC121. The research reported here was undertaken in an effort to assess the potential of optimal measure transport methods as components within the larger scope of a com putational framework envisioned to support research in the radiation biology domain. Within this effort, our interest centered on enabling a unified generic framework where probabilistic modeling, inference, and statistical learning can be carried out for a wide range of data distributions. As described next in Section 1 (and in more detail in our original publication), optimal measure transport offers the possibility of such unified approach.

97 MATHEMATICS AND COMPUTING↗

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

Applications of LIF to Document Natural Variability of Chlorophyll Content and Cu Uptake in Moss

Chlorophyll has long been used as a natural indicator of plant health and photosynthetic efficiency. Laser-induced fluorescence (LIF) is an emerging technique for understanding broad spectrum organic processes and has more recently been used to monitor chlorophyll response in plants. Previous work has focused on developing a LIF technique for imaging moss mats to identify metal contamination with the current focus shifting toward application to moss fronds and aiding sample collection for chemical analysis. Two laser systems (CoCoBi a Nd:YGa pulsed laser system and Chl-SL with two blue continuous semiconductor diodes) were used to collect images of moss fronds exposed to increasing levels of Cu (1, 10, and 100 nmol/cm 2 ) using a CMOS camera. The best methods for the preprocessing of images were conducted before the analysis of fluorescence signatures were compared to a control. The Chl-SL system performed better than the CoCoBi, with dynamic time warping (DTW) proving the most effective for image analysis. Manual thresholding to remove lower decimal code values improved the data distributions and proved whether using one or two fronds in an image was more advantageous. A higher DTW difference from the control correlated to lower chlorophyll a/b ratios and a higher metal content, indicating that LIF, with the aid of image processing, can be an effective technique for identifying Cu contamination shortly after an event.

59 BASIC BIOLOGICAL SCIENCES↗

Portable Software Environment for Ultrahigh-Resolution ELM Development on GPUs

This paper presents our endeavors in developing the large-scale, ultra-high-resolution E3SM Land Model (uELM), specifically designed for exascale computers furnished with accelerators such as Nvidia GPUs. The uELM is a sophisticated code that substantially relies on High-Performance Computing (HPC) environments, necessitating particular machine and software configurations. To facilitate community-based uELM developments employing GPUs, we have created a portable, standalone software environment preconfigured with uELM input datasets, simulation cases, and source code. This environment, utilizing Docker, encompasses all essential code, libraries, and system software for uELM development on GPUs. It also features a functional unit test framework and an offline model testbed for comprehensive numerical experiments. From a technical perspective, the paper discusses GPU-ready container generations, uELM code management, and input data distribution across computational platforms. Lastly, the paper demonstrates the use of environment for functional unit testing, end-to-end simulation on CPUs and GPUs, and collaborative code development.

E3SM Land Model↗

Dependence of CCN closure relationship with organic fraction from two airborne field campaigns over mid-latitude land and ocean

This study investigates the relationship between measured and calculated cloud condensation nuclei (CCN) number concentration and its dependence with organic fraction utilizing aircraft observations from The Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA, 2017–2018) and The Holistic Interactions of Shallow Clouds, Aerosols, and Land Ecosystems (HI-SCALE, 2016) campaigns, which represent midlatitude marine and continental environments, respectively. For the ACE-ENA marine region, aerosol and CCN concentrations were significantly higher in summer than in winter, whereas at continental site for HI-SCALE, aerosol and CCN concentrations showed no pronounced differences between spring and autumn. Using aerosol chemical composition and number size distribution data, CCN concentrations at various supersaturations are calculated based on Köhler theory and then compared with observations from CCN counter. The results show that CCN closure performs well at both sites with a slight overestimation, with mean closure ratio (CR) of 1.13 and 1.17, respectively. Further investigation reveals that CR at lower supersaturation perform better than that at higher supersaturation. The dependence of CR on organic mass fraction (MForg) varies by environment: for marine aerosols, CR decreases with increasing organic fraction at lower supersaturations, whereas continental aerosols exhibit a consistent overestimation, with CR decreasing as organic fraction increases at higher supersaturations. This study provides key insights into CCN characteristics over midlatitude marine and continental environments, emphasizing the necessity of incorporating size-resolved chemical composition and mixing states into future model parameterizations, and contributing to a better understanding of aerosol–cloud interactions.

ACE-ENA field campaign↗