Search NASA⌕ Search

SEARCH · Search NASA

Results for “computational complexity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Behavior and mechanisms of Doppler wind lidar error in complex terrain: stable flow case study at Perdigão

A numerical experiment is carried out investigating the magnitude of biases in ground-based lidar measurements in complex flow conditions. Biases assessed include those arising from flow curvature and from the interaction of turbulence with the wind field reconstruction (WFR) algorithms used by a WindCube lidars and anemometers. RANS-CFD and WRF-LES simulations were performed for the Perdig˜ao Field Experiment site for a range of atmospheric conditions. Virtual anemometer and lidar data were generated for four locations: two near exposed ridge tops and two in low-speed regions in the valley. The LES data at these four locations show that the scalar inflation terms (the relation between scalar and vector averaged wind speed) for virtual lidar and virtual cups agree very well with predictions using perturbation theory. While the lidar errors vary greatly with location and height, the contribution from the flow curvature tends to be larger than the differences arising from scalar inflation. For one lidar/mast pair near the ridge top, comparisons between simulations and measurements are carried out for a resonant mountain wave event on June 14th, 2017, and for the whole duration of the Perdigão campaign for winds perpendicular to the ridges. The lidar error during the mountain wave, a period of strong stability and low inversion height, is significantly larger than the campaign average. The sensitivity of the lidar error to atmospheric stability is confirmed by the RANS simulations, which suggests strong sensitivity of flow curvature error to stability conditions and to the shape of the wind speed profile near the top of the boundary layer.

17 WIND ENERGY↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Light-ray wave functions and integrability

Using integrability, we construct (to leading order in perturbation theory) the explicit form of twist-three light-ray operators in planar $\mathcal{N}$ = 4 SYM. This construction allows us to directly compute analytically continued CFT data at complex spin. We derive analytically the “magic” decoupling zeroes previously observed numerically. Using the Baxter equation, we also show that certain Regge trajectories merge together into a single unifying Riemann surface. Perhaps more surprisingly, we find that this unification of Regge trajectories is not unique. If we organize twist-three operators differently into what we call “cousin trajectories” we find infinitely more possible continuations. We speculate about which of these remarkable features of twist-three operators might generalize to other operators, other regimes and other theories.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GrainPaint: A multi-scale diffusion-based generative model for microstructure reconstruction of large-scale objects

Simulation-based approaches to microstructure generation can suffer from a variety of limitations, such as high memory usage, long computational times, and difficulties in generating complex geometries. Generative machine learning models present a way around these issues, but they have previously been limited by the fixed size of their generation area. Here, we present a new microstructure generation methodology leveraging advances in inpainting using denoising diffusion models to overcome this generation area limitation. We show that microstructures generated with the presented methodology are statistically similar to grain structures generated with a kinetic Monte Carlo simulator, SPPARKS.

36 MATERIALS SCIENCE↗

Towards robust surrogate models: Benchmarking machine learning approaches to expediting phase field simulations of brittle fracture

Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.

Benchmark dataset↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Deciphering the Cofilin Oligomers via Intermolecular Disulfide Bond Formation: A Coarse-Grained Molecular Dynamics Approach to Understanding Cofilin’s Regulation on Actin Filaments

Cofilin, a key actin-binding protein, orchestrates the dynamics of the actomyosin network through its actin-severing activity and by promoting the recycling of actin monomers. Recent experimental work suggests that cofilin also forms functionally distinct oligomers through thiol post-translational modification (PTM) that encourages actin nucleation and assembly. Despite these advances, the structural conformations of cofilin oligomers that modulate actin activity remain elusive because there are combinatorial ways to oxidize thiols in cysteines to form disulfide bonds rapidly. This study employs molecular dynamics simulations to investigate human cofilin 1 as a case study for exploring cofilin dimers via disulfide bond formation. Using the free energy profiling, our simulations unveil a range of probable cofilin dimer structures not represented in current Protein Data Bank entries. These candidate dimers are characterized by their distinct population distributions and relative free energies. Of particular note is a dimer featuring an interface between cysteines 139 and 147 residues, which demonstrates stable free energy characteristics and intriguingly symmetrical geometry. In contrast, the experimentally proposed dimer structure exhibits a less stable free energy profile. Here, we also evaluate frustration quantification based on the energy landscape theory in the protein-protein interactions at the dimer interfaces. Notably, the 39-39 dimer configuration emerges as a promising candidate for forming cofilin tetramers, as substantiated by frustration analysis. Additionally, docking simulations with actin filaments further evaluate the stability of these cofilin dimer-actin complexes. Our findings thus offer a computational framework for understanding the role of thiol post-translational modification cofilin proteins in regulating oligomerization, and the subsequent cofilin-mediated actin dynamics in the actomyosin network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Density Functional Tight Binding Insights into Plasmonic Silver–Platinum Nanoparticles and Alloys for Enhanced Photocatalysis

Developing accurate and efficient Slater-Koster (SK) tight-binding parameter sets is essential for quantum plasmonic studies of alloyed metal nanoparticles, as conventional time dependent density functional theory (TD-DFT) calculations are computationally prohibitive for larger clusters. In this work, we develop and validate density functional tight binding (DFTB) parameter sets for both ground state (GS-SK) and excited state (ES-SK) calculations to study the structural, electronic, and optical properties of silver (Ag), platinum (Pt), and Ag–Pt nanoalloys. Our investigation of the ground state properties demonstrates that the GS-SK parameters enable DFTB to closely reproduce the electronic structures of platinum clusters with diverse sizes and geometries – showing qualitative agreement with DFT for density of states (DOS) profiles and energy levels. The ES-SK parameters accurately describe excited-state properties compared to TD-DFT reference calculations, including the broad, featureless absorption profiles of Pt that are dominated by interband transitions. Using the ES-SK parameters within a real-time TD-DFTB framework, we compute size-dependent optical absorption spectra of Ag, Pt and Ag-Pt nanocubes containing up to 1099 atoms (size ∼4.18 nm). A detailed study of Ag–Pt and Pt-Ag core–shell nanoparticles shows quenching of the Ag plasmon resonance even at monolayer coverage for Ag-Pt, but not for Pt-Ag. We also show how to define submonolayer Ag-core Pt-shell cubic structures that have similar optical properties to those generated experimentally for much larger particles, which offers potential for describing plasmon-enhanced photocatalysis. Collectively, the GS-SK and ES-SK parameter sets provide an accurate, computationally efficient approach for modeling the complex optical and electronic behavior of noble–transition metal nanostructures and their alloys.

SPR↗

Efficient learning of accurate surrogates for simulations of complex systems

Machine learning methods are increasingly deployed to construct surrogate models for complex physical systems at a reduced computational cost. However, the predictive capability of these surrogates degrades in the presence of noisy, sparse or dynamic data. Here, we introduce an online learning method empowered by optimizer-driven sampling that has two advantages over current approaches: it ensures that all local extrema (including endpoints) of the model response surface are included in the training data, and it employs a continuous validation and update process in which surrogates undergo retraining when their performance falls below a validity threshold. We find, using benchmark functions, that optimizer-directed sampling generally outperforms traditional sampling methods in terms of accuracy around local extrema even when the scoring metric is biased towards assessing overall accuracy. Finally, the application to dense nuclear matter demonstrates that highly accurate surrogates for a nuclear equation-of-state model can be reliably autogenerated from expensive calculations using few model evaluations.

79 ASTRONOMY AND ASTROPHYSICS↗

A Second Moment Method for k -Eigenvalue Acceleration with Continuous Diffusion and Discontinuous Transport Discretizations

The second moment method is a linear acceleration technique that couples the transport equation to a diffusion equation with transport-dependent additive closures. The resulting low-order diffusion equation can be discretized independent of the transport discretization, unlike diffusion synthetic acceleration, and is symmetric positive definite, unlike quasidiffusion. While this method has been shown to be comparable to quasidiffusion in iterative performance for fixed source and time-dependent problems, it is largely unexplored as an eigenvalue problem acceleration scheme due to the belief that the resulting inhomogeneous source makes the problem ill posed. Recently, a preliminary feasibility study was performed on the second moment method for eigenvalue problems. The results suggested comparable performance to quasidiffusion and more robust performance than diffusion synthetic acceleration. This work extends the initial study to more realistic reactor problems using state-of-the-art discretization techniques. Finally, the results in this paper show that the second moment method is more computationally efficient than its alternatives on complex reactor problems with unstructured meshes.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Accelerating iterative ptychography with an integrated neural network

Electron ptychography is a powerful and versatile tool for high-resolution and dose-efficient imaging. Iterative reconstruction algorithms are powerful but also computationally expensive due to their relative complexity and the many hyperparameters that must be optimised. Gradient descent-based iterative ptychography is a popular method, but it may converge slowly when reconstructing low spatial frequencies. Here, in this work, we present a method for accelerating a gradient descent-based iterative reconstruction algorithm by training a neural network (NN) that is applied in the reconstruction loop. The NN works in Fourier space and selectively boosts low spatial frequencies, thus enabling faster convergence in a manner similar to accelerated gradient descent algorithms. We discuss the difficulties that arise when incorporating a NN into an iterative reconstruction algorithm and show how they can be overcome with iterative training. We apply our method to simulated and experimental data of gold nanoparticles on amorphous carbon and show that we can significantly speed up ptychographic reconstruction of the nanoparticles.

4DSTEM↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

Metalens formed by structured arrays of atomic emitters

Abstract Arrays of atomic emitters have proven to be a promising platform to manipulate and engineer optical properties, due to their efficient cooperative response to near‐resonant light. Here, we theoretically investigate their use as an efficient metalens. We show that, by spatially tailoring the (subwavelength) lattice constants of three consecutive two‐dimensional arrays of identical atomic emitters, one can realize a large transmission coefficient with arbitrary position‐dependent phase shift, whose robustness against losses is enhanced by the collective response. To characterize the efficiency of this atomic metalens, we perform large‐scale numerical simulations involving a substantial number of atoms (N∼ 5 × 10 5 ) that is considerably larger than comparable works. Our results suggest that low‐loss, robust optical devices with complex functionalities, ranging from metasurfaces to computer‐generated holograms, could be potentially assembled from properly engineered arrays of atomic emitters.

Materials Science↗

Cavitand-Mediated Photodimerization of Chalcones: The Effect of Supramolecular Influences and Temperature on Reaction Selectivity

The photocycloaddition (PCA) of chalcones represents an important reaction pathway for accessing substituted cyclobutanes, which is a molecular framework with utility in synthetic chemistry, materials science, and medicine. In the past, our group has demonstrated the utility of the large cavity of γ-CD as a container for encapsulating two photo reactants for directing the PCA of several classes of aryl alkenes with high stereo- and regioselectivity: the cavitand-mediated photodimerization (CMP) approach. The CMP of chalcones reported in this work further demonstrates the effectiveness of this approach as high yields of dimers were observed in the photoreactions, while they were non-reactive in the solid state and yielded only the isomerization product in homogeneous media. The γ-CD CMP of chalcones yielded predominantly dimerized products in very good to high yields (>70%), composed of a mixture of three dimers in different proportions with syn HH as the major product. Computational analysis of the ground state complex structures revealed a strong correlation between the stability of the complex and predominance of the stereoisomer in the mixture. Further insights were deduced from temperature-dependence studies, which showed a shift in dimer selectivity tending towards a single stereoisomer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gas-Phase Coordination of Phosphine-Chalcogenides to the Uranyl Cation

It is generally believed that actinides exhibit an increased covalency in bonding to soft-donor atoms such as nitrogen and sulfur compared to the lanthanides. The explanation is that the greater spatial extent of the 5f orbitals in the actinides compared to the 4f orbitals in the lanthanides accounts for this behavior. However, recent computational studies on lanthanide and actinide complexes with sulfur-donor atom ligands suggest that while calculated metrics of bond covalency could increase for actinide complexes with sulfur ligands, the overall stability of the complexes decreased. While many studies of actinide covalency were conducted in the condensed phases, this work investigates the intrinsic interaction between actinides and soft-donor ligands compared with oxygen in the gas phase.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gas-Phase Coordination of Phosphine-Chalcogenides to the Uranyl Cation

It is generally believed that actinides exhibit an increased covalency in bonding to soft-donor atoms such as nitrogen and sulfur compared to the lanthanides. The explanation is that the greater spatial extent of the 5f orbitals in the actinides compared to the 4f orbitals in the lanthanides accounts for this behavior. However, recent computational studies on lanthanide and actinide complexes with sulfur-donor atom ligands suggest that while calculated metrics of bond covalency could increase for actinide complexes with sulfur ligands, the overall stability of the complexes decreased. While many studies of actinide covalency were conducted in the condensed phases, this work investigates the intrinsic interaction between actinides and soft-donor ligands compared with oxygen in the gas phase.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗