Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Stress intensity factor models using mechanics-guided decomposition and symbolic regression

The finite element method can be used to compute accurate stress intensity factors (SIFs) for cracks with complex geometries and boundary conditions. In contrast, handbook solutions act as surrogate SIF models that provide significantly faster evaluation times. However, the development of conventional surrogate SIF models relies on manual development based on low-order parameterizations. This limits surrogate model accuracy and generalizability. Here, in this paper, we develop a framework for the automated development of mechanics-guided handbook SIF solutions by using interpretable machine learning via genetic programming for symbolic regression (GPSR). Formalizing the mechanics-based approach of Raju and Newman, SIF training data is decomposed into multiple subsets. This decomposition enables parallel GPSR model development of subfunctions, each of which accounts for specific geometrical corrections with respect to a known analytical model. Using this mechanics-based approach with GPSR allows for equations to be learned with improved accuracy and reduced complexity relative to the Raju Newman equations while maintaining the inherent interpretability of mathematical expressions. In this paper, we present equations that match the complexity of the Raju Newman equations while having reduced error, as well as equations with similar errors and reduced complexity.

42 ENGINEERING↗

Designed Spin‐Texture‐Lattice to Control Anisotropic Magnon Transport in Antiferromagnets

Abstract Spin waves in magnetic materials are promising information carriers for future computing technologies due to their ultra‐low energy dissipation and long coherence length. Antiferromagnets are strong candidate materials due, in part, to their stability to external fields and larger group velocities. Multiferroic antiferromagnets, such as BiFeO 3 (BFO), have an additional degree of freedom stemming from magnetoelectric coupling, allowing for control of the magnetic structure, and thus spin waves, with the electric field. Unfortunately, spin‐wave propagation in BFO is not well understood due to the complexity of the magnetic structure. In this work, long‐range spin transport is explored within an epitaxially engineered, electrically tunable, 1D magnonic crystal. A striking anisotropy is discovered in the spin transport parallel and perpendicular to the 1D crystal axis. Multiscale theory and simulation suggest that this preferential magnon conduction emerges from a combination of a population imbalance in its dispersion, as well as anisotropic structural scattering. This work provides a pathway to electrically reconfigurable magnonic crystals in antiferromagnets.

36 MATERIALS SCIENCE↗

A coarse-grained model of clay colloidal aggregation and consolidation with explicit representation of the electrical double layer

The aggregation of clay minerals in liquid water exemplifies colloidal self-assembly in nature. These negatively charged aluminosilicate platelets interact through multiple mechanisms with different sensitivities to particle shape, surface charge, aqueous chemistry, and interparticle distance and exhibit complex aggregation structures. Experiments have difficulty resolving the associated colloidal assemblages at the scale of individual particles. Conversely, all-atom molecular dynamics (MD) simulations provide detailed insight on clay colloidal interaction mechanisms, but they are limited to systems containing a few particles. We develop a new coarse-grained (CG) model capable of representing assemblages of hundreds of clay particles with accuracy approaching that of MD simulations, at a fraction of the computational cost. Our CG model is parameterized based on MD simulations of a pair of smectite clay particles in liquid water. A distinctive feature of our model is that it explicitly represents the electrical double layer (EDL), i.e., the cloud of charge-compensating cations that surrounds the clay particles. Our model captures the simultaneous importance of long-range colloidal interactions (i.e., interactions consistent with simplified analytical models, already included in extant clay CG models) and short-range interactions such as ion correlation and surface and ion hydration effects. The resulting simulations correctly predict, at low solid-water ratios, the existence of ordered arrangements of parallel particles separated by water films with a thickness up to ~10 nm and, at high solid-water ratios, the coexistence of crystalline and osmotic swelling states, in agreement with experimental observations.

54 ENVIRONMENTAL SCIENCES↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.5)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Coarse-grained simulation of colloidal self-assembly, cation exchange, and rheology in Na/Ca smectite clay gels

Knowledge Gap: The aggregation of clay minerals—layered silicate nanoparticles—strongly impacts fluid flow, solute migration, and solid mechanics in soils, sediments, and sedimentary rocks. Experimental and computational characterization of clay aggregation is inhibited by the delicate water-mediated nature of clay colloidal interactions and by the range of spatial scales involved, from 1 nm thick platelets to flocs with dimensions up to micrometers or more. Simulations: Using a new coarse-grained molecular dynamics (CGMD) approach, we predicted the microstructure, dynamics, and rheology of hydrated smectite (more precisely, montmorillonite) clay gels containing up to 2,000 clay platelets on length scales up to 0.1 μm. Further, simulations investigated the impact of simulation time, platelet diameters (6 to 25nm), and the ratio of Na to Ca exchangeable cations on the assembly of tactoids (i.e., stacks of parallel clay platelets) and larger aggregates (i.e., assemblages of tactoids). We analyzed structural features including tactoid size and size distribution, basal spacing, counterion distribution in the electrical double layer, clay association modes, and the rheological properties of smectite gels. Findings: Our results demonstrate new potential to characterize and understand clay aggregation in dilute suspensions and gels on a scale of thousands of particles with explicit representation of counterion clouds and with accuracy approaching that of all-atom molecular dynamics (MD) simulations. For example, our simulations predict the strong impact of Na/Ca ratio on clay tactoid formation and the shear-thinning rheology of clay gels.

42 ENGINEERING↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

In situ multi-tier auto-ignition detection applied to dual-fuel combustion simulations

Here we use an anomaly detection methodology that is centered on analyzing fourth-order joint moments (co-kurtosis), particularly focusing on its application in auto-ignition of combustion problems with large numbers of species. Unsupervised anomaly detection is challenging to generalize across problem types and domains. A recent technique, centered on analyzing information in the fourth-order joint moment co-kurtosis, has shown promise, especially for high-dimensional scientific data. In this work we present developments to the co-kurtosis based anomaly detection method needed to make it effective and scalable for large-scale distributed scientific data, such as those generated by massively parallel simulations. An in situ co-kurtosis algorithm is employed as the anomaly detection method for identifying ignition kernels in simulations of turbulent combustion. Here, we extend an existing methodology which identifies regions of the domain where anomalies are present, and add another tier of anomaly detection where the individual samples contributing to the anomaly are identified. We apply this algorithm on-the-fly to a variety of turbulent reacting flow problems and compare it to the widely used (but significantly more expensive) chemical explosive mode analysis (CEMA). We demonstrate the ability of the method to detect and identify the onset of low and high temperature ignition which can be used for computational steering, as chemical and combustion anomalies occur intermittently at spatio-temporal locations unknown a priori. Finally, we apply our lightweight in situ algorithm to an exascale high-fidelity simulation with a total of 2.4 Trillion degrees of freedom, performed using an adaptive mesh refinement solver. Furthermore, through a scalability analysis, we show that the relative computational cost of this in-situ anomaly detection algorithm compared to an iteration of the reacting flow solver is negligible.

97 MATHEMATICS AND COMPUTING↗

NEAMS Technical Area Support in MOOSE

The MOOSE framework is a foundational capability used by the NEAMS program to create over 15 different simulation tools for advanced nuclear reactors. Due to this ubiquity, improvements to the framework in support of modeling and simulation goals are critical to the program. These improvements can take many forms including optimization, improved user experience, streamlined application programming interfaces (APIs), parallelism, and other new capabilities. The work transcribed in this report was conducted in direct support of the simulation tools and has already been deployed. The capabilities outlined in this report include enabling selective polynomial basis refinement, implementing a custom convergence system, building a scalable preconditioner for saddle-point problems, and much more.

97 MATHEMATICS AND COMPUTING↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Peeling-ballooning modes in spherical tokamaks: Multi-branch instabilities and effects beyond ideal MHD

A number of important physics effects on the stability of relatively high-n (n is the toroidal mode number) peeling-ballooning modes (PBMs) are investigated based on an equilibrium reconstructed from a NSTX discharge, utilizing extended magnetohydrodynamic (MHD) eigenvalue solvers. For a given toroidal mode number n, multiple branches of instabilities are computed, with the total number of unstable branches roughly linearly scaling with n. Most of the unstable branches are located in the plasma core region, but edge-localized branches, i.e., PBMs, are also identified at higher n-numbers. For the single-fluid-wise most unstable PBM with n = 19⁠, stabilizing/destabilizing effects due to various physics beyond ideal MHD are systematically investigated. Plasma toroidal flow is found to be weakly stabilizing. Local flow shear is generally stabilizing as well, with the degree of stabilization depending on the initial growth rate (without flow shear) of the mode. The plasma resistivity can strongly destabilize the PBM within the single-fluid framework. Anisotropic thermal transport, strong parallel sound wave damping, as well as two-fluid effects are all stabilizing to the mode. In particular, diamagnetic stabilization (within the two-fluid model) is found to be very strong for this mode.

Linear stability analysis↗

Exploring Anomalous Photoelectron Angular Distributions in the Photoelectron Spectra of Gd 3 O 3 – : Study of Gd 3 O 2 – and Gd 3 O 3 – Using Photoelectron Spectroscopy and Density Functional Theory Calculations

Anion photoelectron (PE) spectra of lanthanide oxide clusters obtained previously have exhibited anomalous photoelectron angular distributions which were attributed to strong PE–valence electron (PEVE) interactions. Here, to further explore this effect, we have obtained the PE spectra of Gd 3 O 2 – and Gd 3 O 3 – , two clusters that have similarly complex electronic structures but contrasting symmetries. The spectra exhibit manifolds of detachment transitions at similar binding energies in a 0.5 eV window of energy. The electron affinity of Gd 3 O 2 is measured to be 1.29 ± 0.05 eV, and that of Gd 3 O 3 is 1.31 ± 0.05 eV. As seen in previous studies on lanthanide oxide cluster anions in lower than conventional oxidation states, transitions in spectra obtained lower photon energies are more congested than those obtained with higher photon energy, a signature of strong PEVE interactions. While the detachment transitions have predominantly parallel photoelectron angular distributions (PAD), the PAD varies across the manifold of transitions in the PE spectrum of Gd 3 O 3 – in a way that suggests four different subgroups of transitions. Results of calculations on Gd 3 O 2 – suggest kite or V-shape structures with antiferromagnetic coupling between one of the 4f 7 subshells with the two others. Calculations on Gd 3 O 3 – more definitively point to ring structures with a nearly isoenergetic ferromagnetically coupled high spin (24-tet) state and a dectet state in which one of the 4f 7 subshells is antiferromagnetically coupled with the other two. Taking these results as qualitative, we propose that strong mixing between the unperturbed states predicted computationally leads to overlapping transitions with different PADs.

anions↗

Snow ALbedo eVOlution (SALVO) Campaign Broadband Albedo from April - June, 2024 in Utqiagivk, AK level a1

A field-portable broadband (285 – 2800 nm) albedometer was used to make spatially distributed albedo measurements on tundra and sea ice surfaces. The albedometer consists of paired upward-looking and downward-looking pyranometers, which were both connected to a data logger. The instrument was mounted approximately 1 m above the surface using a tripod and was placed on a 1.4 m-long boom to minimize the impacts of shading from the operator and to observe surfaces undisturbed by footprints (see Appendix for photos of measurement setup and uncertainty assessment). Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1.2 m south of the line. On the operator’s end of the boom, there was a bubble level that was aligned with the bubble level on the upward-looking pyranometer. To take a measurement, the operator first relocated the tripod to the measurement location, then leveled the instrument and held it level for at least twice the pyranometers’ response time (5 or 15 seconds, see below), and finally depressed a trigger on the data logger. The data logger recorded the instantaneous voltage on both pyranometers, the measurement number, and the time. The data logger also converted the voltages to irradiances, and from these computed the ratio (outgoing/incoming) for albedo, which could be checked in the field. The operator recorded in a field notebook the measurement number that corresponded with the locations on the line and any pertinent notes (e.g., invalid measurements). With this setup, a trained operator could measure a 200-m albedo line (41 measurements) in approximately 30 minutes. Measurements were made within 3 hours of solar noon. The data logger had sufficient storage capacity to record all measurements from the campaign, but data were downloaded to a computer after each measurement day.

54 ENVIRONMENTAL SCIENCES↗

Deployment of BISON models of fuel restructuring at high burnup and related fission gas behavior in UO 2

This milestone report details the advancements made in fiscal year 2024 under the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program to improve the modeling of fission gas behavior in high burnup UO 2 nuclear fuel in the BISON fuel performance code. As nuclear fuel is pushed to higher burnups, significant microstructural changes occur within the fuel, including the formation of a high burnup structure (HBS) on the pellet rim and a dark zone deeper within the pellet. These regions, characterized by subgrain formation and increased pore densities, have critical implications for fission gas behavior and release, which are not well understood. The modeling capabilities in BISON did not adequately predict these phenomena, leading to an underestimation of fuel restructuring and - potentially - of fission gas release. To address these gaps, this milestone focused on three key objectives: (1) reviewing and assessing Sifgrs's capabilities for low burnup fuel, on which high burnup capabilities rely, (2) validating and expanding HBS fission gas modeling capabilities, including investigating mechanisms for fission gas release from HBS, and (3) expanding Sifgrs to enable modeling of dark zone formation and its effects on fission gas behavior. These objectives were achieved and are described herein. The achievements of this NEAMS milestone are significant for the industry's goal of burnup extension. The improved predictive modeling capabilities for both low- and high-burnup conditions enhance our understanding of fuel performance under both normal operations and transient scenarios. Although goals were reached, future work is necessary to validate these models against experimental data and quantify their accuracy in different conditions. In parallel, mechanistic modeling efforts should continue to extend and refine these capabilities to increase accuracy while reducing reliance on empirical models. This will ensure robust performance across a broader range of conditions.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Analysis of Infrastructures for Processing Plastic Waste using Pyrolysis-Based Chemical Upcycling Pathways

Modern mechanical recycling infrastructure for plastic is capable of processing only a small subset of waste plastics, reinforcing the need for parallel disposal methods such as landfilling and incineration. Emerging pyrolysis-based chemical technologies can "upcycle" plastic waste into high-value polymer and chemical products and process a broader range of waste plastics. In this work, we study the economic and environmental benefits of deploying an upcycling infrastructure in the continental United States for producing low-density polyethylene (LDPE) and polypropylene (PP) from post-consumer mixed plastic waste. Our analysis aims to determine the market size that the infrastructure can create, the degree of circularity that it can achieve, the prices for waste and derived products it can propagate, and the environmental benefits of diverting plastic waste from landfill and incineration facilities it can produce. We apply a computational framework that integrates techno-economic analysis, life cycle assessment, and value chain optimization. Our results demonstrate that the infrastructure generates an economy of nearly 20 billion USD and positive prices for plastic waste, opening opportunities for compensation to residents who provide plastic waste. Our analysis also indicates that the infrastructure can achieve a plastic-to-plastic degree of circularity of 34% and remains viable under various external factors (including technology efficiencies, capital investment budgets, and polymer market values). Finally, we present significant environmental benefits of upcycling over alternative landfill and incineration waste disposal methods, and comment on ongoing work expanding our modeling methodology to other chemical upcycling pathway case studies, including hydroformylation of specific plastics to chemicals.

Interdisciplinary↗

Analysis and modeling of tungsten emission and net erosion in the DIII-D divertor using updated atomic data

Tungsten (W) is one of the leading candidate materials for plasma-facing components. However, its main drawback is its high radiative efficiency; if W penetrates the plasma, it can lead to core degradation or even collapse. Since eroded tungsten tends to ionize in the sheath and redeposit promptly, the net erosion flux that escapes prompt redeposition can differ significantly from the gross erosion. This work presents a modeling framework to estimate net erosion and photon emission from W coatings exposed to the lower divertor of DIII-D using the DiMES material exposure probe. The approach couples RustBCA for sputtering yields with a Monte Carlo transport code (LPTMC) that models redeposition and W emission. Computation is carried out with new atomic data, based on R-matrix and Mons calculations, leading to lower ionization probabilities and a twofold increase in net erosion estimates compared to calculations done with OPEN-ADAS atomic data. The model results are benchmarked against experimental measurements, showing quantitative agreement for erosion, although the trends in W emission are reproduced only qualitatively. The model is also used to assess whether W II emission can serve as a direct measurement of the net erosion of W in the lower divertor of DIII-D. Simulations show that this is not valid if the electron pressure is above ~120 Pa or if the toroidal length of the eroded material is smaller than the parallel-to-B distance traveled by impurity ions before steady-state conditions are reached. Finally, simulations suggest that when W is sputtered by carbon ions with high impact energies (≳300 eV) in DIII-D, W net erosion scales with W gross erosion and can be numerically approximated using W I flux alone as input.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Supercharging simulation-based inference for Bayesian optimal experimental design

Abstract Bayesian optimal experimental design (BOED) seeks to maximize the expected information gain (EIG) of experiments. This requires a likelihood estimate, which in many settings is intractable. Simulation-based inference (SBI) provides powerful tools for this regime. However, existing work explicitly connecting SBI and BOED is restricted to a single contrastive EIG bound. We show that the EIG admits multiple formulations which can directly leverage modern SBI density estimators, encompassing neural posterior, likelihood, and ratio estimation. Building on this perspective, we define a novel EIG estimator using neural likelihood estimation. Further, we identify optimization as a key bottleneck of gradient based EIG maximization and show that a simple multi-start parallel gradient ascent procedure can substantially improve reliability and performance. With these innovations, our SBI-based BOED methods are able to match or outperform by up to 22% existing state-of-the-art approaches across standard BOED benchmarks.

97 MATHEMATICS AND COMPUTING↗

Safe Reinforcement Learning-Based Transient Stability Control for Islanded Microgrids With Topology Reconfiguration

This paper proposes a safe reinforcement learning (RL)-based transient stability emergency control (TSEC) method for islanded microgrids. RL requires extensive interaction with the environment to learn control strategies, hence, a data-driven approach is used as a substitute for time-consuming time-domain simulation calculations. Deep sigma point processes (DSPP), which is a Gaussian process model, is utilized to predict the normal distribution of transient stability of microgrids and to construct a transient stability chance constraint. Reward-constrained policy optimization (RCPO) can simultaneously achieve objective prediction, policy learning, and constraint cost coefficient update across multiple timescales. RCPO interacts with the DSPP-based microgrid environment through a multi-process parallel manner, greatly increasing the training speed. Case studies on a real islanded microgrid demonstrate that the proposed method can efficiently and quickly obtain the optimal emergency control strategy while adhering to all hard constraints.

14 SOLAR ENERGY↗