Search NASA⌕ Search

SEARCH · Search NASA

Results for “adaptive sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Persistent Sampling: Enhancing the Efficiency of Sequential Monte Carlo

Sequential Monte Carlo (SMC) samplers are powerful tools for Bayesian inference but suffer from high computational costs due to their reliance on large particle ensembles for accurate estimates. We introduce persistent sampling (PS), an extension of SMC that systematically retains and reuses particles from all prior iterations to construct a growing, weighted ensemble. By leveraging multiple importance sampling and resampling from a mixture of historical distributions, PS mitigates the need for excessively large particle counts, directly addressing key limitations of SMC such as particle impoverishment and mode collapse. Crucially, PS achieves this without additional likelihood evaluations-weights for persistent particles are computed using cached likelihood values. This framework not only yields more accurate posterior approximations but also produces marginal likelihood estimates with significantly lower variance, enhancing reliability in model comparison. Furthermore, the persistent ensemble enables efficient adaptation of transition kernels by leveraging a larger, decorrelated particle pool. Experiments on high-dimensional Gaussian mixtures, hierarchical models, and non-convex targets demonstrate that PS consistently outperforms standard SMC and related variants, including recycled and waste-free SMC, achieving substantial reductions in mean squared error for posterior expectations and evidence estimates, all at reduced computational cost. PS thus establishes itself as a robust, scalable, and efficient alternative for complex Bayesian inference tasks.

Karamanis, Minas↗

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

Simulant Development of Potential 200 West Area Waste Feeds

Preliminary planning for retrieval, qualification, and pretreatment of waste in Hanford’s 200 West Area (200W) has begun as part of the West Area Risk Management project. Experimental studies to technically mature pretreatment process operations will likely be needed because of the uniqueness of 200W waste. Pacific Northwest National Laboratory formulated five simulants to represent 200W-qualified feed based on the preliminary flowsheet provided by Washington River Protection Solutions, LLC. The simulant recipes were devised using applicable historical information as a reference point to support the use of the flowsheet waste vectors, which were combined into five distinct groups. These five groups formed the basis for the liquid composition targets that were adapted into recipes using charged-balanced salt species. The liquid phase recipes were batched in 1-L quantities and analyzed at Pacific Northwest National Laboratory. Once confirmed to be stable, the liquid solutions were tested for compatibility with candidate solid components. Specific solid components were recommended based on cross-examining the proposed solid phases in the flowsheet with relevant data from the literature. Mixtures of solid components were added to aliquots of the liquid batches and sub-sampled to measure particle size distribution. The measured distribution was compared to independently created benchmark distributions appropriate for each simulant. This process was iterated until a solid phase composition that resulted in a representative particle size distribution was found. After the final compositions were confirmed, a suite of chemical and physical characterization data was collected. This report describes the simulant basis, formulation methodology, laboratory measurements, and data collected for the recipes recommended to represent 200W waste feeds.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Virulence and Genetic Diversity of Puccinia spp., Causal Agents of Rust on Switchgrass (Panicum virgatum L.) in the USA

Switchgrass (Panicum virgatum L.) is an important cellulosic biofuel grass native to North America. Rust, caused by Puccinia spp. is the most predominant disease of switchgrass and has the potential to impact biomass conversion. In this study, virulence patterns were determined on a set of 38 switchgrass genotypes for 14 single-spore rust isolates from 14 field samples collected in seven states. Single nucleotide polymorphism (SNP) variation was also assessed in 720 sequenced cloned amplicons representing 654 base pairs of the elongation factor 1-α gene from the field samples. Five major haplotypes were identified differing by 11 out of the 39 SNP positions identified. STRUCTURE, Principal Coordinate Analysis, and phylogenetic analyses divided the rust population into two genetic clusters. Virginia and Georgia had the highest and lowest rust genetic diversity, respectively. Only nine accessions showed a differential disease response between the 14 isolates, allowing the identification of eight races, differing by 1–3 virulence factors. Overall, the results suggested clonal reproduction of the pathogen and a North–South differentiation via local adaptation. However, similar haplotypes and races were also recovered from several states, suggesting migration events, and highlighting the need to further investigate the switchgrass rust population structure and evolution in the USA.

Bahri, Bochra A. (ORCID:0000000159055880)↗

Dynamic Control of Sodium Cold Trap Purification Temperature Using LSTM System Identification

This study investigates the dynamic regulation of the sodium cold trap purification temperature at Argonne National Laboratory’s liquid sodium test facility, employing long short-term memory (LSTM) system identification techniques. The investigation introduces an innovative hybrid approach by integrating model predictive control (MPC) based on first principles dynamic models with a multi-step time–frequency LSTM model in predicting the temperature profiles of a sodium cold trap purification system. The long short-term memory–model predictive controller (LSTM-MPC) model employs a sliding window scheme to gather training samples for multi-step prediction, leveraging historical data to construct predictive models that capture the non-linearities of the complex system dynamics without explicitly modeling the underlying physical processes. The performance of the LSTM-MPC and MPC were evaluated through simulation experiments, where both models were assessed on their capacity to maintain the cold trap temperature within predefined set-points while minimizing deviations and overshoots. Results obtained show how the data-driven LSTM-MPC model demonstrates stability and adaptability. In contrast, the traditional MPC model exhibits irregularities, particularly evident as overshoots around set-point limits, which can potentially compromise its effectiveness over long prediction time intervals. The findings obtained offer valuable insights into integrating data-driven techniques for enhancing real-time monitoring systems.

LSTM-MPC↗

MUSE adaptive-optics spectroscopy confirms dual active galactic nuclei and strongly lensed systems at sub-arcsec separation

The novel Gaia multi peak (GMP) technique has proven to be able to successfully select dual and lensed active galactic nuclei (AGN) candidates at sub-arcsecond separations. Both populations are important because dual AGN represent one of the central, still largely untested, predictions of ΛCDM cosmology, and compact lensed AGN allow us to probe the central regions of the lensing galaxies. In this work, we present high-spatial-resolution spectroscopy of 12 GMP-selected systems. We used the adaptive-optics assisted integral-field spectrograph MUSE at the VLT to resolve each system and investigate the nature of each component. All targets show the presence of two components confirming the GMP selection. We classify 4 targets as dual AGN, 3 as lensed quasar candidates, and 5 as a chance alignment of a star and an AGN. With separations ranging from 0.30″ to 0.86″, these dual and lensed systems are among the most compact systems discovered to date at z > 0.5. This is the largest sample of distant dual AGN with sub-arcsecond separations ever presented in a single paper.

Scialpi, M. (ORCID:0009000651004986)↗

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

3D Geologic Framework Modelling of the Los Alamos National Laboratory Site and Pajarito Plateau: Integrating a realistic 3D fault network and modelling subsurface relationships in a sparsely sampled and complex geologic region

The subsurface geology beneath the Pajarito Plateau is critical to understanding the seismic hazard of the Pajarito Fault System, yet our understanding of this geology is relatively poor. While previous 3D geologic framework models of the area have been created for the purposes of understanding hydrogeologic flow, they are inadequate for the purposes of understanding the Pajarito Fault System. The specific challenges of using oil and gas software for this purpose include: (1) the geologic complexities resulting from volcanism and tectonism; (2) a need for a high level of stratigraphic detail over a large area; (3) a near complete lack of seismic data; and (4) sparse wellbore data. Presented here is a workflow that handles these challenges of adapting commercially available software used by the oil and gas industries to this seismic hazard problem.

58 GEOSCIENCES↗

FunDiff: diffusion models over function spaces for physics-informed generative modeling

Recent advances in generative modeling-particularly diffusion models and flow matching-have been widely used for synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. To address this, we introduce FunDiff, an efficient and robust framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generates continuous functions that can be evaluated at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, demonstrating that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We further demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results indicate that our method can generate physically consistent samples with high fidelity to the target distribution, and exhibit robustness to noisy and low-resolution data.

Wang, Sifan [Yale University, New Haven, CT (Unite↗

Instrumentation and methods for efficient time-resolved X-ray crystallography of biomolecular systems with sub-10 ms time resolution

Time-resolved X-ray crystallography has great promise to illuminate structure–function relations and key steps of enzymatic reactions with atomic resolution. The dominant methods for chemically-initiated reactions require complex instrumentation at the X-ray beamline, significant effort to operate and maintain this instrumentation, and enormous numbers (∼10 5 –10 9 ) of crystals per time point. We describe instrumentation and methods that enable high-throughput time-resolved study of biomolecular systems using standard crystallography sample supports and mail-in X-ray data collection at standard high-throughput cryocrystallography synchrotron beamlines. The instrumentation allows rapid reaction initiation by mixing of crystals and substrate/ligand solution, rapid capture of structural states via thermal quenching with no pre-cooling perturbations, and yields time resolutions in the single-millisecond range, comparable to the best achieved by any non-photo-initiated method in both crystallography and cryo-electron microscopy. Our approach to reaction initiation has the advantages of simplicity, robustness, low cost, adaptability to diverse ligand solutions and small minimum volume requirements, making it well suited to routine laboratory use and to high-throughput screening. We report the detailed characterization of instrument performance, present structures of binding of N -acetylglucosamine to lysozyme at time points from 8 ms to 2 s determined using only one crystal per time point, and discuss additional improvements that will push time resolution toward 1 ms.

Indergaard, John A. (ORCID:000000022367699X)↗

Fabrication of porous transport electrodes: Development of quantitative approach for quality control

This work focuses on porous transport electrodes (PTEs), which integrate the anodic catalyst with the adjacent Ti porous transport layer (PTL). Challenges in catalyst deposition on PTLs, particularly at low loadings, motivated this study to evaluate various fabrication methods and characterization approaches. This work investigated Pt-treated PTLs coated with Ir-based catalysts using several common methods, including airbrush coating, rod coating, ultrasonic spray coating, electrodeposition, and sputter deposition, with catalyst loadings ranging from 2.9 to 0.1 mg/cm 2 , providing the opportunity for comparisons across a large set of samples produced by different methods. Two widely accessible characterization techniques: X-ray computed tomography (XCT) and scanning electron microscopy energy dispersive X-ray spectroscopy (SEM-EDS) were explored. Initial evaluation of selected samples with XCT provided qualitative insights into catalyst distribution, however comprehensive quantitative analysis was limited. SEM-EDS enabled detailed information on the catalyst distribution both qualitatively and quantitatively using two metrics. Atomic and surface area % ratios of Pt:Ir and Ti:Ir revealed trends in catalyst loading and losses into the PTL pores, as well as evaluating the homogeneity of catalyst coatings. The analysis demonstrated that ultrasonic spray coating, electrodeposition, and sputter coating produced the most homogeneous coatings, with minimal catalyst losses observed for electrodeposition and sputter coating. By adapting common techniques with novel, standardized methodologies, this work establishes a universally applicable framework for cross-study comparison of PTEs. The SEM-EDS approach provides a practical, accessible tool for PTE characterization and contributes a reference dataset supporting both research development and rapid quality control.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE↗

Development of Automated Atom Probe Tomography capability to study the influence of applied voltage and laser power on the final apparent composition of the analyzed specimen

This study presents the development and implementation of an autonomous Bayesian optimization (BO) framework for controlling and optimizing experimental parameters in Atom Probe Tomography (APT). Using commercial silicon needle samples as a benchmark system, we demonstrate that BO can efficiently navigate the complex parameter space of voltage and laser power to achieve target charge state ratios (specifically Si + /(Si + +Si 2+ )) with minimal experimental evaluations. Our implementation integrates Gaussian Process modeling with the CAMECA atom probe control framework, enabling autonomous adjustment of experimental conditions in real-time. Results show that the algorithm successfully converges to target ratios under different scenarios: maintaining a reference ratio, increasing the ratio (favoring Si 1+ ), and decreasing the ratio (favoring Si 2+ ). The system adapts to specimen evolution during analysis, compensating for changes in apex geometry while maintaining optimization targets. This work establishes a proof of concept for AI-driven optimization in APT, addressing the traditional challenges of manual parameter tuning and paving the way for applications to more complex materials where compositional accuracy is critical.

36 MATERIALS SCIENCE↗

Computationally efficient models for aqueous organic redox flow batteries

The rising usage of intermittent energy has garnered the need for large scale energy storage systems. Redox flow batteries (RFB) based energy storage system shows promising potential. Numerical simulations and machine learning approaches have been widely used to study RFB performance. The development of autonomous material discovery framework and digital twin of energy storage system usually needs to query cell performance through fast response models. In this study, two computationally efficient models are introduced: a physics-based analytical flow battery model (EZBattery), and a machine learning operator model (Deep Operator Network, denoted by DeepONet). Both models can provide cell performance near instantly, and prediction accuracy was systematically examined on an application of evaluating the performances of a 780 cm 2 aqueous organic redox flow battery (AORFB), using potential anolyte candidates in dihydroxyphenazine (DHP)-based family of organic materials. A validated computationally expansive 3-dimensional multi-physics finite element model by COMSOL was used as the ground truth and provided the training data set for the DeepONet. 1280 samples were generated with 10 properties to mimic the different possible anolyte candidates, and the cell performances were evaluated under 10 different combined operating conditions. The accuracy comparisons for the two computationally efficient models show that both models can provide comparable accuracy in predicting cell charging/discharging voltage curves. DeepONet can provide slightly higher overall accuracy than EZBattery with faster calculation speed, but highly relies on the training dataset. EZBattery does not need a training dataset and can provide interpretable physics-based explanations of the results, while being more flexible to adjust to adapt any different cell designs, flow battery architectures, and electrolyte materials.

Analytical model↗

A new cw-NMR Q-meter for dynamically polarized targets for particle physics

Polarized solid targets produced via Dynamic Nuclear Polarization rely on Continuous-Wave Nuclear Magnetism Resonance measurements to accurately determine the degree of polarization of bulk samples polarized to nearly 100%. Since the late 1970's phase sensitive detection methods have been utilized to observe the magnetization of a sample as a small change in inductance under RF excitation near the Larmor frequency of the nuclear species of interest, using a device known as a Q-meter. Liverpool Q-meters, produced in the UK in the 80's and 90's, have been the workhorse devices for these targets for decades, however their age and scarcity has meant new systems are needed. In conclusion, we describe a Q-meter system designed and built at Jefferson Lab in the Liverpool style to have comparable electronic performance with several improvements to update and adapt the devices for modern use.

Dynamic nuclear polarization↗

The evolution of analytical techniques for multiplex analysis of protein biomarkers

Introduction: The landscape of biomarker development has evolved with advanced analytical technologies, particularly affinity- and mass spectrometry-based techniques. These advancements have deepened our understanding of disease mechanisms, enabling the development of precise diagnostic tools and personalized medicine. Protein biomarkers, which play pivotal roles in biological processes, have become invaluable in diagnosing and monitoring diseases, aided by their presence in various biological samples and the availability of established detection methods. Areas covered: This review covers the role of protein biomarkers in clinical practice, the development and dimensionality of protein biomarkers, advancements in detection technologies, a comparison of these technologies, and future directions in biomarker discovery and disease mechanism elucidation. Expert opinion: Advances in biomarker technologies have the potential to transform diagnostics and personalized treatment but face challenges such as high costs and technical complexity. Enhancing reproducibility and integrating multi-omics approaches may offer better insights. In conclusion, the field should evolve toward high-throughput, automated methods, continuously adapting research, and clinical practices.

59 BASIC BIOLOGICAL SCIENCES↗