Search NASA⌕ Search

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan↗

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

Dynamics modeling of molten salt reactor with reduced and expanded representations of delayed neutron precursors

Molten salt reactors (MSRs) present unique challenges in dynamic behavior due to the mobility of their fuel. In these reactors, delayed neutron precursors (DNPs) drift with the fuel circulation through the primary loop. As a result, a fraction of DNPs decays outside the core, effectively reducing the available delayed neutron population for reactivity control. Consequently, precise modeling of the distribution and behavior of DNPs is critical for accurate reactor dynamics simulations. In this study, the System Dynamics Analysis Tool (SDAT) was used to simulate a thermal-spectrum MSR under steady-state conditions and following transients. The effects of using reduced and expanded representations of DNPs with fewer or more groups than the conventional 6-group model were investigated. Their impact on the simulated distribution of precursors in the primary loop, reactivity loss value, and reactor response to transients was analyzed. Simulation results showed that reduced models lead to the loss of the actual DNPs distribution data, resulting in less accurate estimates of reactivity loss. Reactor power predictions using these reduced models showed significant deviations compared to those using the conventional 6-group model in transient simulations. Expanded models offered a more accurate representation of the distribution of DNPs and reactivity loss estimates. Reactor power predictions using expanded models showed minimal deviation from the conventional 6-group model during the simulated transients.

analysis↗

SITCOMTN-154: Initial studies of photometric redshifts with LSSTComCam from DP1

This technote holds reports based on the first analyses of the Data Preview 1 (DP1) data by the Science Unit for photometric redshifts. Although photometric redshifts are not an official DP1 data product, the "Photo-z Science Unit" generated photo-z estimates for every galaxy in DP1 using the available multi-band imaging on a best-effort basis. This work included developing training and test datasets by matching DP1 data to high-quality reference redshifts obtained with spectroscopy, Grism data, and multi-band photometry. The Science Unit used the RAIL software package to make photometric redshift estimates using eight different algorithms, developed simple scientific performance metrics, used those metrics to explore how the performance of the algorithms varied with configuration changes, derived more optimized configurations of the algorithms and tested the performance of those configurations. This work, the resulting data products and expected data distribution mechanism are all described there.

79 ASTRONOMY AND ASTROPHYSICS↗

FEDERATED LEARNING ON STOCHASTIC NEURAL NETWORKS

Federated learning is a machine learning paradigm that leverages edge computing on client devices to optimize models while maintaining user privacy by ensuring that local data remain on the device. However, since all data are collected by clients, federated learning is susceptible to latent noise in local datasets. Factors such as limited measurement capabilities or human errors may introduce inaccuracies in client data. To address this challenge, we propose the use of a stochastic neural network as the local model within the federated learning framework. Stochastic neural networks not only facilitate the estimation of the true underlying states of the data but also enable the quantification of latent noise. We refer to our federated learning approach, which incorporates stochastic neural networks as local models, as federated stochastic neural networks. In this work we will present numerical experiments demonstrating the performance and effectiveness of our method, particularly in handling nonindependent and identically distributed data.

97 MATHEMATICS AND COMPUTING↗

Establishing Pb-203 production from electrodeposited Tl targets at Brookhaven National Laboratory

Background: Promising developments in Pb-212 radiopharmaceutical therapies have increased demand for Pb-203 diagnostic agents. Building on previous work from various isotope production facilities, this study optimized Pb-203 production from electrodeposited Tl targets at Brookhaven National Laboratory (BNL). The additional supply of Pb-203 may help meet growing preclinical and clinical demands. Results: Two Tl targets were irradiated at the Brookhaven Linac Isotope Producer facility with 30 ± 1 MeV protons, measured using previously published cross section data. Distribution coefficients for Pb Resin in acetate media were investigated for both Na + and K + cations, where potassium acetate was ~ 4 times more effective at stripping Pb from the Pb Resin. The Tl electrodeposition was optimized to deposit 350 mg of Tl (~ 60 mg/cm 2 ) on Au backing in under 6 h. The proposed separation process was completed in < 1.5 h and achieved > 98% and 92 ± 3% recovery of Tl and Pb, respectively, with an overall Tl-Pb separation factor of 6 × 10 5 . The experimentally measured half-life of Pb-203 was 52.4 ± 0.7 h, agreeing with 51.93 ± 0.02 h reported by the National Nuclear Data Center. The radioisotopic purity of the Pb fraction at 24 h post end of bombardment (EOB) from a 24 h irradiation was 66% Pb-203, 28% Pb-201, and 6% Pb-200. Following chemical separation, the Pb-203 produced in this work (21 MBq Pb-203 EOB) achieved apparent molar activities of 10 ± 5 and 0.9 ± 0.5 GBq/µmol for [ 203 Pb]Pb-DOTAM and [ 203 Pb]Pb-DO3A, respectively, decay corrected to EOB. Data derived from this work suggests BNL can produce > 10’s GBq Pb-203 with > 99% radiochemical and radioisotopic purity from Tl-205 for worldwide distribution. Conclusions: The production and separation of Pb-203 from natural Tl target material was successfully demonstrated at BNL. Existing methods were adapted and optimized for the facilities at BNL. Results from this work will guide future large-scale Pb-203 production opportunities at BNL for clinical applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

C 12 ( n , n 1 ′ γ ) partial γ -ray cross section measured using the GENESIS array

Improved neutron inelastic scattering cross sections have repeatedly been identified as a top priority nuclear data need, important for basic science and a range of applications in nuclear energy, stockpile stewardship, and proliferation detection. For the C 12 ( n , n ′ γ ) reaction in particular, recent measurements have unveiled some structural discrepancies, demonstrating incongruities among themselves and in relation to the ENDF/B-VIII.0 nuclear data evaluation. To help resolve these disagreements, a measurement was performed at the 88-Inch Cyclotron at Lawrence Berkeley National Laboratory using a broad-spectrum neutron beam and a 99.8% pure natural carbon target. The Gamma Energy Neutron Energy Spectrometer for Inelastic Scattering (GENESIS) was employed to measure energy-differential γ -ray emission spectra as a function of incident neutron energy in the energy range of 5.5 to 16.7 MeV. The C 12 partial γ -ray cross sections were extracted at 63 ∘ , 122 . 5 ∘ , and 150 ∘ with respect to the incoming neutron beam and integrated using angular distribution data available in the literature. The data show agreement with a recent literature measurement and evaluation from 11 to 15 MeV, but indicate a larger cross section for incident neutron energies between 5.5 and 8.5 MeV. The measured relative angular distributions are also reported and were found to agree with evaluation. Published by the American Physical Society 2025

Gordon, J. M. (ORCID:0009000789886897)↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Utah FORGE: Well 16B(78)-32 Distributed Temperature Sensing Data from April and May 2024

This dataset includes Neubrex Energy Services fiber optic distributed temperature sensing (DTS) data from well 16B(78)-32 during stimulation and circulation, including interaction with well 16A(78)-32, during April and May 2024. The DTS data are stored in HDF5 file format and are accompanied by a PowerPoint report on the study. All times in this dataset are in UTC. Depths are in MD relative to Kelly Bushing Height, and temperatures are in degrees Fahrenheit. All DTS measurements were made using a Yokogawa 3000DTSX Distributed Temperature Sensing Interrogator Unit, with a spatial sampling interval of 3.28 feet and a temporal sampling rate of 129 seconds. The third-party Pressure-Temperature Gauge data should be used with caution after April 20, 2024, as its performance is not considered reliable beyond this date.

15 GEOTHERMAL ENERGY↗

Physicochemical and Molecular Insights into the Boundary Layer and Free Troposphere Aerosol Interactions over the Southern Great Plains

Ambient aerosols’ vertical profiles are critical for evaluating the role of aerosols in atmospheric chemistry and radiative transfer, but limited data on these profiles hinders our ability to fully assess their impact on the Earth's radiative balance. Here, in this study, we investigated the size-, time-, and altitude resolved composition of individual particles and bulk molecular composition of particle samples collected by an uncrewed aerial system–ArcticShark over the Southern Great Plains. Single particle microanalysis shows that, the free tropospheric (FT) samples are dominated (56-66%) by carbonaceous sulfate particles, while boundary layer (BL) samples are dominated (57-74%) by carbonaceous particles. Back trajectory simulations suggest that FT particles are likely influenced by long-range transport and have undergone aqueous-phase processing. Conversely, in-situ size distribution data shows evidence of particle growth in the upper BL and just below the FT. This observation may indicate vertical transport of particles from an elevated aerosol layer in the FT, possibly linked to a new particle formation event. This observation is further supported by high resolution molecular composition data, which reveals particle volatility increasing with increasing size, which aligns with the growth event. This study aids in fundamental understanding of the compositional and molecular specificity of vertically resolved organic aerosols to provide insights into particle size evolution for future atmospheric models.

ArcticShark↗

SIDDA: SInkhorn Dynamic Domain Adaptation

Modern neural networks (NNs) often do not generalize well in the presence of a "covariate shift"; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more domain-invariant features. Domain adaptation (DA) methods include a range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SIDDA, an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, and real astronomical observations. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with equivariant neural networks (ENNs). We find that SIDDA enhances the generalization capabilities of NNs, achieving up to a ≈40% improvement in classification accuracy on unlabeled target data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, we find that SIDDA enhances model calibration on both source and target data--achieving over an order of magnitude improvement in the ECE and Brier score. SIDDA's versatility, combined with its automated approach to domain alignment, has the potential to advance multi-dataset studies by enabling the development of highly generalizable models.

Pandya, Sneh [Northeastern U.]↗

Systematic engineering for production of anti-aging sunscreen compound in Pseudomonas putida

Sunscreen has been used for thousands of years to protect skin from ultraviolet radiation. However, the use of modern commercial sunscreen containing oxybenzone, ZnO, and TiO 2 has raised concerns due to their negative effects on human health and the environment. In this study, we aim to establish an efficient microbial platform for production of shinorine, a UV light absorbing compound with anti-aging properties. First, we methodically selected an appropriate host for shinorine production by analyzing central carbon flux distribution data from prior studies alongside predictions from genome-scale metabolic models (GEMs). We enhanced shinorine productivity through CRISPRi-mediated downregulation and utilized shotgun proteomics to pinpoint potential competing pathways. Simultaneously, we improved the shinorine biosynthetic pathway by refining its design, optimizing promoter usage, and altering the strength of ribosome binding sites. Finally, we conducted amino acid feeding experiments under various conditions to identify the key limiting factors in shinorine production. The study combines meta-analysis of 13 C-metabolic flux analysis, GEMs, synthetic biology, CRISPRi-mediated gene downregulation, and omics analysis to improve shinorine production, demonstrating the potential of Pseudomonas putida KT2440 as platform for shinorine production.

59 BASIC BIOLOGICAL SCIENCES↗

GPU-Accelerated Solution of the Bethe–Salpeter Equation for Large and Heterogeneous Systems

We present a massively parallel GPU-accelerated implementation of the Bethe–Salpeter equation (BSE) for the calculation of the vertical excitation energies (VEEs) and optical absorption spectra of condensed and molecular systems, starting from single-particle eigenvalues and eigenvectors obtained with density functional theory. The algorithms adopted here circumvent the slowly converging sums over empty and occupied states and the inversion of large dielectric matrices through a density matrix perturbation theory approach and a low-rank decomposition of the screened Coulomb interaction, respectively. Further computational savings are achieved by exploiting the nearsightedness of the density matrix of semiconductors and insulators to reduce the number of screened Coulomb integrals. We scale our calculations to thousands of GPUs with a hierarchical loop and data distribution strategy. The efficacy of our method is demonstrated by computing the VEEs of several spin defects in wide-band-gap materials, showing that supercells with up to 1000 atoms are necessary to obtain converged results. We discuss the validity of the common approximation that solves the BSE with truncated sums over empty and occupied states. In conclusion, we then apply our GW-BSE implementation to a diamond lattice with 1727 atoms to study the symmetry breaking of triplet states caused by the interaction of a point defect with an extended line defect.

Absorption spectra↗

Regularization via f -Divergence: An Application to Multi-Oxide Spectroscopic Analysis

In this paper, we explore the application of convolutional neural networks (CNNs) for predicting the chemical composition of complex geologic samples in a simulated Martian atmospheric environment. Specifically, we aim to characterize oxide weight percentages (wt.%) of rock samples analyzed by remote Laser-Induced Breakdown Spectroscopy (LIBS), framing the problem as a multi-target regression task . Neural networks trained on LIBS spectra are prone to overfitting due to high spectral complexity, limited labeled data, and measurement noise. While regularization is critical for improving generalization, common methods (e.g., ℓ 2 regularization) impose constraints not directly tied to data distribution properties. We propose a novel regularization method based on a specific ƒ-divergence induced by a graph-based estimator, designed to constrain the distributional discrepancy between predictions and targets. This regularizer serves a dual purpose: (a) mitigating overfitting by enforcing a constraint on the distributional difference between predictions and noisy targets, and (b) acting as an auxiliary loss that penalizes large divergences. To enable backpropagation, we develop a differentiable approximation of this particular ƒ-divergence, making the method feasible for neural networks. Experiments on ChemCam and SuperCam LIBS calibration spectra show that mathematical equation-divergence regularization outperforms or matches standard regularization methods (ℓ 1 , ℓ 2 , dropout) and the classical baseline, partial least squares (PLS). Combining ƒ-divergence regularization with standard regularization yields further performance gains, indicating that distributional regularization is useful in this context giving a promising direction for robust model training in planetary science applications. Source code is publicly available at Klein and Li (2025), https://doi.org/10.11578/dc.20250530.7.

58 GEOSCIENCES↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Measurement of the small-scale 3D Lyman- α forest power spectrum

Small-scale correlations measured in the Lyman-α (Lyα) forest encode information about the intergalactic medium and the primordial matter power spectrum. In this article, we present and implement a simple method to measure the 3-dimensional power spectrum, P 3D , of the Lyα forest at wavenumbers k corresponding to small, ~ Mpc scales. In order to estimate P 3D from sparsely and unevenly distributed data samples, we rely on averaging 1-dimensional Fourier Transforms, as previously carried out to estimate the 1-dimensional power spectrum of the Lyα forest, P 1D . Further, this methodology exhibits a very low computational cost. We confirm the validity of this approach through its application to Nyx cosmological hydrodynamical simulations. Subsequently, we apply our method to the eBOSS DR16 Lyα forest sample, providing as a proof of principle, a first P 3D measurement averaged over two redshift bins z = 2.2 and z = 2.4. This work highlights the potential for forthcoming P 3D measurements, from upcoming large spectroscopic surveys, to untangle degeneracies in the cosmological interpretation of P 1D .

79 ASTRONOMY AND ASTROPHYSICS↗