Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computational Complexity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Experiments to validate Thermodynamics and Transport models of Strongly Coupled Dusty Plasma Matter (Final Technical Report for DE-SC0023416)

The goal of this two-year grant is to provide access to the PI to dusty plasma experimental facilities at the DOE-funded Collaborative Research Facility Magnetized Plasma Research Laboratory, Auburn University to become a user of that facility, to obtain experimental data to support another ongoing grant DE-SC0021146 (an Early Career Award to the PI that is focused on modeling of dusty plasma thermodynamics and transport processes), to generate experimental data for funding proposals, and to provide exposure to University of Memphis students to advanced experimental techniques. The following technical accomplishments were made: 1. Development of a novel Bidirectional Electrode Control Arms Assembly (BECAA) for producing perfect 2D grain layers for complex plasma experimentation. BECAA uses movable electrode arms to tilt or move the electrode in a RF discharge from outside the chamber, allowing for the manipulation of grain clouds without needing to change the plasma parameters or gas pressure. This work addresses a longstanding gap in the literature for a method to produce clusters of selectable number of grains and that are perfectly two dimensional as opposed to being only quasi-2D. 2. Experimental investigation of the structural properties of finite-N clusters with N=2 to 50. Individual particle behavior in clusters could vary from grain to grain and this study measured systematically produced clusters for two different grain sizes. Analysis (funded by another grant DE-SC0021146) is currently underway to quantify the differences between grains that are found on the surface vs. the interior of clusters, the shell structure, and the decay of correlations in position, velocity, and kinetic energy. 3. An experimental method to measure the structural entropy of clusters was developed by observing the self-induced structural transitions between various possible arrangements. In a series of heating and cooling cycles, the number of times each possible arrangement was attained was experimentally observed and used to compute the probability of existence of that arrangement, and subsequently the configurational entropy of the cluster. Analysis (funded by another grant DE-SC0021146) is currently underway to produce the entropy of clusters as a function of the number of grains and use the same to compute thermodynamic state variables for 2D complex plasma/grain clusters. 4. A preliminary experimental study of multibody collisions between N grains (N=2 – 10) was conducted. The clusters were produced using the BECAA technique and velocities were imparted to the grains using manipulation laser pulses. Analysis (funded by another grant DE-SC0021146) is currently underway to develop a theoretical framework to describe multibody collisions analogous to classical two-body interactions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Operational Analytics Studies for ATLAS Distributed Computing: Data Popularity Forecast and Utilization of the WLCG Centers

Operational analytics is the direction of research related to the analysis of the current state of computing processes and the prediction of future states in order to anticipate imbalances and take timely measures to stabilize a complex system. There are two relevant areas in ATLAS Distributed Computing that are currently the focus of studies: user physics analysis including the forecast of popularity of data samples among users, and evaluating WLCG centers for their readiness to process user analysis payloads. Studying these areas is challenging due to the complexity involved, as it requires a comprehensive understanding of numerous boundary conditions typically found in large-scale distributed computing infrastructures. Forecasts of data popularity are problematic without the categorization of user tasks by their types (data transformation or physics analysis), which do not always appear on the surface but may induce noise, which introduces significant distortions for predictive analysis. Evaluating the WLCG resources by their analysis workloads is also a challenging task as it is necessary to find a balance between the workload of the resource, its performance, the waiting time for jobs on it, as well as the volume of jobs that it processes. This is especially difficult in a heterogeneous computing environment, where legacy resources are used along with modern high-performance machines. We will look at these areas of research in detail and discuss what tools and methods are used in our work, demonstrating results already obtained.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Behavior and mechanisms of Doppler wind lidar error in complex terrain: stable flow case study at Perdigão

A numerical experiment is carried out investigating the magnitude of biases in ground-based lidar measurements in complex flow conditions. Biases assessed include those arising from flow curvature and from the interaction of turbulence with the wind field reconstruction (WFR) algorithms used by a WindCube lidars and anemometers. RANS-CFD and WRF-LES simulations were performed for the Perdig˜ao Field Experiment site for a range of atmospheric conditions. Virtual anemometer and lidar data were generated for four locations: two near exposed ridge tops and two in low-speed regions in the valley. The LES data at these four locations show that the scalar inflation terms (the relation between scalar and vector averaged wind speed) for virtual lidar and virtual cups agree very well with predictions using perturbation theory. While the lidar errors vary greatly with location and height, the contribution from the flow curvature tends to be larger than the differences arising from scalar inflation. For one lidar/mast pair near the ridge top, comparisons between simulations and measurements are carried out for a resonant mountain wave event on June 14th, 2017, and for the whole duration of the Perdigão campaign for winds perpendicular to the ridges. The lidar error during the mountain wave, a period of strong stability and low inversion height, is significantly larger than the campaign average. The sensitivity of the lidar error to atmospheric stability is confirmed by the RANS simulations, which suggests strong sensitivity of flow curvature error to stability conditions and to the shape of the wind speed profile near the top of the boundary layer.

17 WIND ENERGY↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Light-ray wave functions and integrability

Using integrability, we construct (to leading order in perturbation theory) the explicit form of twist-three light-ray operators in planar $\mathcal{N}$ = 4 SYM. This construction allows us to directly compute analytically continued CFT data at complex spin. We derive analytically the “magic” decoupling zeroes previously observed numerically. Using the Baxter equation, we also show that certain Regge trajectories merge together into a single unifying Riemann surface. Perhaps more surprisingly, we find that this unification of Regge trajectories is not unique. If we organize twist-three operators differently into what we call “cousin trajectories” we find infinitely more possible continuations. We speculate about which of these remarkable features of twist-three operators might generalize to other operators, other regimes and other theories.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GrainPaint: A multi-scale diffusion-based generative model for microstructure reconstruction of large-scale objects

Simulation-based approaches to microstructure generation can suffer from a variety of limitations, such as high memory usage, long computational times, and difficulties in generating complex geometries. Generative machine learning models present a way around these issues, but they have previously been limited by the fixed size of their generation area. Here, we present a new microstructure generation methodology leveraging advances in inpainting using denoising diffusion models to overcome this generation area limitation. We show that microstructures generated with the presented methodology are statistically similar to grain structures generated with a kinetic Monte Carlo simulator, SPPARKS.

36 MATERIALS SCIENCE↗

Towards robust surrogate models: Benchmarking machine learning approaches to expediting phase field simulations of brittle fracture

Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.

Benchmark dataset↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Deciphering the Cofilin Oligomers via Intermolecular Disulfide Bond Formation: A Coarse-Grained Molecular Dynamics Approach to Understanding Cofilin’s Regulation on Actin Filaments

Cofilin, a key actin-binding protein, orchestrates the dynamics of the actomyosin network through its actin-severing activity and by promoting the recycling of actin monomers. Recent experimental work suggests that cofilin also forms functionally distinct oligomers through thiol post-translational modification (PTM) that encourages actin nucleation and assembly. Despite these advances, the structural conformations of cofilin oligomers that modulate actin activity remain elusive because there are combinatorial ways to oxidize thiols in cysteines to form disulfide bonds rapidly. This study employs molecular dynamics simulations to investigate human cofilin 1 as a case study for exploring cofilin dimers via disulfide bond formation. Using the free energy profiling, our simulations unveil a range of probable cofilin dimer structures not represented in current Protein Data Bank entries. These candidate dimers are characterized by their distinct population distributions and relative free energies. Of particular note is a dimer featuring an interface between cysteines 139 and 147 residues, which demonstrates stable free energy characteristics and intriguingly symmetrical geometry. In contrast, the experimentally proposed dimer structure exhibits a less stable free energy profile. Here, we also evaluate frustration quantification based on the energy landscape theory in the protein-protein interactions at the dimer interfaces. Notably, the 39-39 dimer configuration emerges as a promising candidate for forming cofilin tetramers, as substantiated by frustration analysis. Additionally, docking simulations with actin filaments further evaluate the stability of these cofilin dimer-actin complexes. Our findings thus offer a computational framework for understanding the role of thiol post-translational modification cofilin proteins in regulating oligomerization, and the subsequent cofilin-mediated actin dynamics in the actomyosin network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Density Functional Tight Binding Insights into Plasmonic Silver–Platinum Nanoparticles and Alloys for Enhanced Photocatalysis

Developing accurate and efficient Slater-Koster (SK) tight-binding parameter sets is essential for quantum plasmonic studies of alloyed metal nanoparticles, as conventional time dependent density functional theory (TD-DFT) calculations are computationally prohibitive for larger clusters. In this work, we develop and validate density functional tight binding (DFTB) parameter sets for both ground state (GS-SK) and excited state (ES-SK) calculations to study the structural, electronic, and optical properties of silver (Ag), platinum (Pt), and Ag–Pt nanoalloys. Our investigation of the ground state properties demonstrates that the GS-SK parameters enable DFTB to closely reproduce the electronic structures of platinum clusters with diverse sizes and geometries – showing qualitative agreement with DFT for density of states (DOS) profiles and energy levels. The ES-SK parameters accurately describe excited-state properties compared to TD-DFT reference calculations, including the broad, featureless absorption profiles of Pt that are dominated by interband transitions. Using the ES-SK parameters within a real-time TD-DFTB framework, we compute size-dependent optical absorption spectra of Ag, Pt and Ag-Pt nanocubes containing up to 1099 atoms (size ∼4.18 nm). A detailed study of Ag–Pt and Pt-Ag core–shell nanoparticles shows quenching of the Ag plasmon resonance even at monolayer coverage for Ag-Pt, but not for Pt-Ag. We also show how to define submonolayer Ag-core Pt-shell cubic structures that have similar optical properties to those generated experimentally for much larger particles, which offers potential for describing plasmon-enhanced photocatalysis. Collectively, the GS-SK and ES-SK parameter sets provide an accurate, computationally efficient approach for modeling the complex optical and electronic behavior of noble–transition metal nanostructures and their alloys.

SPR↗

Efficient learning of accurate surrogates for simulations of complex systems

Machine learning methods are increasingly deployed to construct surrogate models for complex physical systems at a reduced computational cost. However, the predictive capability of these surrogates degrades in the presence of noisy, sparse or dynamic data. Here, we introduce an online learning method empowered by optimizer-driven sampling that has two advantages over current approaches: it ensures that all local extrema (including endpoints) of the model response surface are included in the training data, and it employs a continuous validation and update process in which surrogates undergo retraining when their performance falls below a validity threshold. We find, using benchmark functions, that optimizer-directed sampling generally outperforms traditional sampling methods in terms of accuracy around local extrema even when the scoring metric is biased towards assessing overall accuracy. Finally, the application to dense nuclear matter demonstrates that highly accurate surrogates for a nuclear equation-of-state model can be reliably autogenerated from expensive calculations using few model evaluations.

79 ASTRONOMY AND ASTROPHYSICS↗

A Second Moment Method for k -Eigenvalue Acceleration with Continuous Diffusion and Discontinuous Transport Discretizations

The second moment method is a linear acceleration technique that couples the transport equation to a diffusion equation with transport-dependent additive closures. The resulting low-order diffusion equation can be discretized independent of the transport discretization, unlike diffusion synthetic acceleration, and is symmetric positive definite, unlike quasidiffusion. While this method has been shown to be comparable to quasidiffusion in iterative performance for fixed source and time-dependent problems, it is largely unexplored as an eigenvalue problem acceleration scheme due to the belief that the resulting inhomogeneous source makes the problem ill posed. Recently, a preliminary feasibility study was performed on the second moment method for eigenvalue problems. The results suggested comparable performance to quasidiffusion and more robust performance than diffusion synthetic acceleration. This work extends the initial study to more realistic reactor problems using state-of-the-art discretization techniques. Finally, the results in this paper show that the second moment method is more computationally efficient than its alternatives on complex reactor problems with unstructured meshes.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Accelerating iterative ptychography with an integrated neural network

Electron ptychography is a powerful and versatile tool for high-resolution and dose-efficient imaging. Iterative reconstruction algorithms are powerful but also computationally expensive due to their relative complexity and the many hyperparameters that must be optimised. Gradient descent-based iterative ptychography is a popular method, but it may converge slowly when reconstructing low spatial frequencies. Here, in this work, we present a method for accelerating a gradient descent-based iterative reconstruction algorithm by training a neural network (NN) that is applied in the reconstruction loop. The NN works in Fourier space and selectively boosts low spatial frequencies, thus enabling faster convergence in a manner similar to accelerated gradient descent algorithms. We discuss the difficulties that arise when incorporating a NN into an iterative reconstruction algorithm and show how they can be overcome with iterative training. We apply our method to simulated and experimental data of gold nanoparticles on amorphous carbon and show that we can significantly speed up ptychographic reconstruction of the nanoparticles.

4DSTEM↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

Metalens formed by structured arrays of atomic emitters

Abstract Arrays of atomic emitters have proven to be a promising platform to manipulate and engineer optical properties, due to their efficient cooperative response to near‐resonant light. Here, we theoretically investigate their use as an efficient metalens. We show that, by spatially tailoring the (subwavelength) lattice constants of three consecutive two‐dimensional arrays of identical atomic emitters, one can realize a large transmission coefficient with arbitrary position‐dependent phase shift, whose robustness against losses is enhanced by the collective response. To characterize the efficiency of this atomic metalens, we perform large‐scale numerical simulations involving a substantial number of atoms (N∼ 5 × 10 5 ) that is considerably larger than comparable works. Our results suggest that low‐loss, robust optical devices with complex functionalities, ranging from metasurfaces to computer‐generated holograms, could be potentially assembled from properly engineered arrays of atomic emitters.

Materials Science↗

Cavitand-Mediated Photodimerization of Chalcones: The Effect of Supramolecular Influences and Temperature on Reaction Selectivity

The photocycloaddition (PCA) of chalcones represents an important reaction pathway for accessing substituted cyclobutanes, which is a molecular framework with utility in synthetic chemistry, materials science, and medicine. In the past, our group has demonstrated the utility of the large cavity of γ-CD as a container for encapsulating two photo reactants for directing the PCA of several classes of aryl alkenes with high stereo- and regioselectivity: the cavitand-mediated photodimerization (CMP) approach. The CMP of chalcones reported in this work further demonstrates the effectiveness of this approach as high yields of dimers were observed in the photoreactions, while they were non-reactive in the solid state and yielded only the isomerization product in homogeneous media. The γ-CD CMP of chalcones yielded predominantly dimerized products in very good to high yields (>70%), composed of a mixture of three dimers in different proportions with syn HH as the major product. Computational analysis of the ground state complex structures revealed a strong correlation between the stability of the complex and predominance of the stereoisomer in the mixture. Further insights were deduced from temperature-dependence studies, which showed a shift in dimer selectivity tending towards a single stereoisomer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gas-Phase Coordination of Phosphine-Chalcogenides to the Uranyl Cation

It is generally believed that actinides exhibit an increased covalency in bonding to soft-donor atoms such as nitrogen and sulfur compared to the lanthanides. The explanation is that the greater spatial extent of the 5f orbitals in the actinides compared to the 4f orbitals in the lanthanides accounts for this behavior. However, recent computational studies on lanthanide and actinide complexes with sulfur-donor atom ligands suggest that while calculated metrics of bond covalency could increase for actinide complexes with sulfur ligands, the overall stability of the complexes decreased. While many studies of actinide covalency were conducted in the condensed phases, this work investigates the intrinsic interaction between actinides and soft-donor ligands compared with oxygen in the gas phase.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗