Search NASA⌕ Search

SEARCH · Search NASA

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Machine Learning for Predicting Multipactor Susceptibility in Planar RF Structures

Multipactor discharge is a persistent challenge in high-power microwave (HPM) and accelerator systems, where secondary electron avalanches can cause heating, vacuum degradation, and failure. This work presents the first supervised machine learning (ML) framework for multipactor prediction, trained on high-fidelity 3D Particle-in-Cell (PIC) simulation data in planar geometries. The model maps operational, geometric, and material-dependent secondary electron yield (SEY) parameters to the time-averaged electron growth rate, enabling rapid reconstruction of susceptibility charts. Among the models evaluated, tree-based ensemble methods such as Random Forest and Extra Trees demonstrate superior generalization to unseen materials compared to neural networks such as multilayer perceptron (MLP). Performance metrics, including Intersection over Union (IoU), Structural Similarity Index Measure (SSIM), and Pearson correlation, show close agreement with simulation benchmarks. Principal Component Analysis attributes generalization limits to material feature-space disjointedness.

43 PARTICLE ACCELERATORS↗

Random Phase Approximation Correlation Energy Using Real-Space Density Functional Perturbation Theory

We present a real-space method for computing the random phase approximation (RPA) correlation energy within Kohn–Sham density functional theory, leveraging the low-rank nature of the frequency-dependent density response operator. In particular, we employ a cubic-scaling formalism based on density functional perturbation theory that circumvents the calculation of the response function matrix, instead relying on the ability to compute its product with a vector through the solution of the associated Sternheimer linear systems. We develop a large-scale parallel implementation of this formalism using the subspace iteration method in conjunction with the spectral quadrature method while employing the Kronecker product-based method for the application of the Coulomb operator and the conjugate orthogonal conjugate gradient method for the solution of the linear systems. We demonstrate convergence with respect to key parameters and verify the method’s accuracy by comparing with plane-wave results. We show that the framework achieves good strong scaling to many thousands of processors, reducing the time to solution for a lithium hydride system with 128 electrons to around 150 s on 4608 processors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Stochastic Trust-Region Algorithm in Random Subspaces with Convergence and Expected Complexity Analyses

Here, this work proposes a framework for large-scale stochastic derivative-free optimization (DFO) by introducing STARS, a trust-region method based on iterative minimization in random subspaces. This framework is both an algorithmic and theoretical extension of a random subspace derivative-free optimization (RSDFO) framework, and an algorithm for stochastic optimization with random models (STORM). Moreover, like RSDFO, STARS achieves scalability by minimizing interpolation models that approximate the objective in low-dimensional affine subspaces, thus significantly reducing per-iteration costs in terms of function evaluations and yielding strong performance on largescale stochastic DFO problems. The user-determined dimension of these subspaces, when the latter are defined, for example, by the columns of so-called Johnson-Lindenstrauss transforms, turns out to be independent of the dimension of the problem. For convergence purposes, inspired by the analyses of RSDFO and STORM, both a particular quality of the subspace and the accuracies of random function estimates and models are required to hold with sufficiently high, but fixed, probabilities. Using martingale theory under the latter assumptions, an almost sure global convergence of STARS to a first-order stationary point is shown, and the expected number of iterations required to reach a desired first-order accuracy is proved to be similar to that of STORM and other stochastic DFO algorithms, up to constants.

97 MATHEMATICS AND COMPUTING↗

Propagation of partially spatially coherent laser beams in instantaneous Kerr media

The propagation of intense, partially spatially coherent laser beams in a medium with instantaneous third-order susceptibility is studied analytically and numerically. For sufficiently high power relative to that required for nonlinear self-focusing, the propagation initially proceeds in two stages. In the first stage, spatial coherence builds up, and in the second stage, the number of speckles reduces. Once the degree of coherence is sufficiently high, whole-beam self-focusing occurs. The beam power is mostly confined within the initial spot radius. Two analytical approaches for describing the evolution of the beam are presented. The method of moments leads to an analytical solution for the rms spot radius that is in excellent agreement with simulations. This method does not require any knowledge of the field statistics beyond the initial conditions and provides no information about the evolution of the individual speckles. The other approach employs a self-similar solution for the second-order coherence function of the field and assumes that the fourth-order coherence function is factorizable and obeys complex circular Gaussian random statistics. The latter method also leads to an analytical expression for the spot radius, but its predictions for the qualitative evolution of the speckles disagree with wave-optics simulations.

lasers↗

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton↗

Maximum a posteriori Ly α estimator (MAPLE): band power and covariance estimation of the 3D Ly α forest power spectrum

We present a novel maximum a posteriori estimator to jointly estimate band powers and the covariance of the three-dimensional power spectrum (P3D) of Ly $\alpha$ forest flux fluctuations, called MAPLE. Our Wiener-filter based algorithm reconstructs a window-deconvolved P3D in the presence of complex survey geometries typical for Ly $\alpha$ surveys that are sparsely sampled transverse to and densely sampled along the line of sight. We demonstrate our method on idealized Gaussian random fields with two selection functions: (i) a sparse sampling of 30 background sources per square degree designed to emulate the current Dark Energy Spectroscopic Instrument; (ii) a dense sampling of 900 background sources per square degree emulating the upcoming Prime Focus Spectrograph Galaxy Evolution Survey. Our proof-of-principle shows promise, especially since the algorithm can be extended to marginalize jointly over nuisance parameters and contaminants, i.e. offsets introduced by continuum fitting. Our code is implemented in JAX and is publicly available on GitHub.

79 ASTRONOMY AND ASTROPHYSICS↗

Significance of Low‐Velocity Zones on Solute Retention in Rough Fractures

Natural fractures are characterized by high internal heterogeneity. This internal variability is the cause of flow channeling, which in turn leads to contaminant transport taking place primarily along the high-velocity channels. Mass exchange between the high-velocity channels and the low-velocity zones has the potential to enhance contaminant retention, due to solute diffusion into the low-velocity zones and subsequent exposure to additional surface area for diffusion into the bordering rock matrix. Here, we derive a random walk particle tracking method for heterogeneous fractures, which includes an additional term to account for the aperture gradient. The method takes into account advection, diffusion in the fracture and matrix diffusion. The developed numerical framework is applied to assess the effect of low-velocity zones in rough self-affine fractures. The results show that diffusion into low-velocity zones has a visible but modest impact on contaminant retention. The magnitude of this impact does not change considerably, regardless of whether diffusion into the rock matrix is considered in the model, and increases for a decreasing average Péclet number of the fracture.

58 GEOSCIENCES↗

Electric dipole excitations near the neutron separation energies in 96 Mo

Electric dipole strength near the neutron separation energy significantly impacts nuclear structure properties and astrophysical scenarios. These excitations are complex in nature and may involve the so-called pygmy dipole resonance (PDR). Transition densities play a crucial role in understanding the nature of nuclear excited states, including collective excitations, as well as in constructing transition potentials in DWBA or coupled-channels equations. In this work, we focus on electric dipole excitations in spherical molybdenum isotopes, particularly 96 Mo, employing fully consistent Hartree-Fock-Bogoliubov (HFB) and Quasiparticle Random Phase Approximation (QRPA) methods. We analyze the dipole strength near the neutron separation energy, which represents the threshold for neutron capture processes, and examine the isospin characteristics of PDR states through transition density calculations. Examination of proton and neutron transition densities reveals distinctive features of each dipole state, indicating their isoscalar and isovector nature. We observe that the primary component in the enhanced low-energy region exhibits isovector character. The PDR displays a mixture of isoscalar and isovector nature, distinguishing it from the isovector giant dipole resonance (IVGDR). These findings lay the groundwork for future investigations into the role of transition densities in reaction models and for their application to inelastic scattering calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Abstract Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcript’s impacts on plant metabolism has employed data from controlled environments. Here, we generated and analyzed data on the variation in the abundance of 26 metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies (GWAS)) and transcripts (using both transcriptome-wide association studies (TWAS) and an explainable artificial intelligence (AI) approach based on random forest (RF)) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

Mathivanan, Ramesh Kanna↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Frontal Slice Approaches for Tensor Linear Systems

Inspired by the row and column action methods for solving large-scale linear systems, in this work, we explore the use of frontal slices for solving tensor linear systems. In particular, this paper presents a novel approach for using frontal slices of a tensor $\mathcal{A}$ to solve tensor linear systems $\mathcal{A} ∗\mathcal{X} = \mathcal{B}$ where ∗ denotes the $t$-product. In addition, we consider variations of this method, including cyclic, block, and randomized approaches, each designed to optimize performance in different operational contexts. Our primary contribution lies in the development and convergence analysis of these methods. Experimental results on synthetically generated and real-world data, including applications such as image and video deblurring, demonstrate the efficacy of our proposed approaches and validate our theoretical findings.

Luo, Hengrui↗

Diabetes-specific formula with standard of care improves glycemic control, body composition, and cardiometabolic risk factors in overweight and obese adults with type 2 diabetes: results from a randomized controlled trial

Background and aims Medical nutrition therapy is important for diabetes management. This randomized controlled trial investigated the effects of a diabetes-specific formula (DSF) on glycemic control and cardiometabolic risk factors in adults with type 2 diabetes (T2D). Methods Participants ( n = 235) were randomized to either DSF with standard of care (SOC) (DSF group; n = 117) or SOC only (control group; n = 118). The DSF group consumed one or two DSF servings daily as meal replacement or partial meal replacement. The assessments were done at baseline, on day 45, and on day 90. Results There were significant reductions in glycated hemoglobin (−0.44% vs. –0.26%, p = 0.015, at day 45; −0.50% vs. −0.21%, p = 0.002, at day 90) and fasting blood glucose (−0.14 mmol/L vs. +0.32 mmol/L, p = 0.036, at day 90), as well as twofold greater weight loss (−1.30 kg vs. –0.61 kg, p < 0.001, at day 45; −1.74 kg vs. –0.76 kg, p < 0.001, at day 90) in the DSF group compared with the control group. The decrease in percent body fat and increase in percent fat-free mass at day 90 in the DSF group were almost twice that of the control group (1.44% vs. 0.79%, p = 0.047). In addition, the percent change in visceral adipose tissue at day 90 in the DSF group was several-fold lower than in the control group (−6.52% vs. –0.95%, p < 0.001). The DSF group also showed smaller waist and hip circumferences, and lower diastolic blood pressure than the control group (all overall p ≤ 0.045). Conclusion DSF with SOC yielded significantly greater improvements than only SOC in glycemic control, body composition, and cardiometabolic risk factors in adults with T2D.

Tey, Siew Ling↗

Derivative-free stochastic optimization via adaptive sampling strategies

In this paper, we present a novel derivative-free framework for solving unconstrained stochastic optimization problems. Many problems in fields ranging from simulation optimization to reinforcement learning to quantum computing involve settings where only stochastic function values are obtained via a zeroth-order oracle, which has no available gradient information and necessitates the usage of derivative-free optimization methodologies. Our approach includes estimating gradients using stochastic function evaluations and integrating adaptive sampling techniques to control the accuracy in these stochastic approximations. Our framework encapsulates several gradient estimation techniques, including standard finite-difference, Gaussian smoothing, sphere smoothing, randomized coordinate finite-difference, and randomized subspace finite-difference methods. We provide theoretical convergence guarantees for our framework and analyze the worst-case iteration and sample complexities associated with each gradient estimation method. Finally, we demonstrate the empirical performance of the methods on logistic regression and nonlinear least squares problems.

Adaptive sampling↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

CORRLA-RS

The CORRLA-RS package provides a suite of statistical methods for sampling multidimensional distributions and to conduct sensitivity and correlation analysis of large scale data in the Rust programming language. The software provides a unique solution to multidimensional constrained sampling problems utilizing a combination of parallelized Markov Chain Monte Carlo methods and traditional rejection sampling. The sensitivity and correlation analysis methods are backed by a high performance randomized singular value decomposition implementation which enables datasets larger than the random access memory (RAM) size to be analyzed. Additionally, CORRLA-RS implements the active subspace identification method using a KD-Tree and the randomized singular value decomposition acting in concert.

Gurecky, William [Oak Ridge National Laboratory (O↗

Spinbox: tools for many-body quantum systems in a Monte Carlo context

Spinbox is a piece of software that facilitates quantum mechanical calculations relevant to Monte Carlo simulation of atomic nuclei. At the front lines of research on the nuclear many-body problem are a large number of supercomputer-scale simulation codes. These codes produce valuable results but can be hard to understand, especially for those without intimate knowledge of the relevant theoretical methods. Thus, tools that fill pedagogical roles are extremely valuable. Spinbox makes it easy for one to replicate and analyze the computational processes relevant to a Quantum Monte Carlo (QMC) simulation that may be difficult to understand/debug/analyze due to the scale of the corresponding simulation software. Spinbox is written in Python using other state-of-the-art Python modules for numerical calculations. While a number of Python libraries exist that are suited to general quantum many-body calculations, the motivation of Spinbox is quite particular. In Diffusion Monte Carlo methods (DMC, GFMC, AFDMC), the central calculation is the imaginary-time propagation of individual samples of the many-body wavefunction. Although quantum wavefunctions generally must be described by a probability distribution over a basis, DMC imbues particles (within one sample) with classical spatial coordinates. This method is unusual, so other Python packages are typically not set up to do this easily. Furthermore, the software has built-in options for nuclear systems assuming isospin symmetry, which can be set up with other libraries but is a nontrivial process to do so. Features: - numerical representation of samples of the many-body wavefunctions, including tensor-product states (used in AFDMC) - numerical representation of many-body operators, including tensor-product operators: general, spin, imaginary-time propagation, etc. - the correct associated arithmetic and algebra, implemented as class methods - classes for representing realistic nuclear two- and three-body Hamiltonians (e.g. Argonne V18, Illinois NNN) - large-scale parallel integration over random variables, crucial for the AFDMC method My goal is to make this package open source so that anyone may use it and contribute to it, particularly other researchers doing AFDMC calculations

Fox, Jordan↗

Evaluating disease surveillance strategies for early outbreak detection in contact networks with varying community structure

Disease surveillance systems allow public health agencies to respond to emerging diseases before they become widespread. Developing such systems requires identifying optimal ways to monitor in the context of an epidemic outbreak; this problem is known as sensor selection. Contact networks represent the dynamics of interaction in a population and are used to model how a disease spreads in a population and to explore strategies of sensor selection. We evaluated five sensor selection strategies on their ability to provide an early warning of a COVID-like outbreak in synthetic contact networks encapsulated in four network scenarios. Three of these scenarios assessed different aspects of community structure. The fourth scenario employed a contact network representing the population and interactions of 6.8 million people in New York City, constructed from an agent-based simulation using census and transportation data. This scenario exemplifies how sensor selection strategies may perform in a real-world, urban context. Our findings suggest that the choice of the optimal strategy depends heavily on the community structure of the network. Strategies that select highly connected nodes or maximize network coverage are the optimal surveillance strategy for outbreak detection in many network community structures. However, a naive implementation of these strategies may fail to provide an early warning at all—including in the New York City scenario. Moreover, these methods are impractical for real-world use as they require knowledge of the underlying contact network. Instead, a selection strategy that starts with a set of random nodes and then performs a random walk through a chain of neighbors reliably provides early warnings without requiring prior knowledge of the network. We find this method, called “random chain”, to be the most pragmatic for implementation in a real-world disease surveillance context.

60 APPLIED LIFE SCIENCES↗