Search NASA⌕ Search

SEARCH · Search NASA

Results for “Heterogeneous computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Practical Considerations for Understanding Surface Reaction Mechanisms Involved in Heterogeneous Catalysis

Acquiring useful knowledge about the active site(s) of a catalyst, nature of reactant–catalyst interactions, nature of reactive intermediates, rate-determining step, reaction rate orders that affect various process parameters, and reaction mechanism as a whole is exceedingly challenging. This is especially true in the case of heterogeneous catalysts due to the complexity of the nature of surface active sites and their nonstatic behavior. Here, we present our perspective on differentiating between various surface reaction mechanisms in light of pioneering studies by leaders in the field, with the aim of clarifying some of the confusion associated with these complex mechanisms, especially the Eley–Rideal mechanism. Using bibliometric analysis, we identify and discuss the following four reactions that most commonly invoke the Eley– Rideal mechanism: H 2 activation, CO oxidation, esterification of alcohols by acids, and selective catalytic reduction (SCR) of NO x with NH 3 . Our analysis of studies utilizing well-suited experimental and computational methodologies for differentiating surface reaction mechanisms suggests that the above-mentioned four reactions do not occur via the Eley–Rideal mechanism. Instead, each reaction occurs via the Langmuir–Hinshelwood mechanism with nonidealities present. Lastly, we highlight practical considerations regarding select experimental (characterization methods and differential kinetics) and computational modeling that we believe can provide useful insights to accurately discern between the various possible reaction mechanisms in heterogeneous catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ion Traps and Packaging for Heterogenous Integration - Chimera

Microfabricated surface ion traps and silicon-based photonics are critical technologies for scaling quantum systems. Current ion trap architectures face scalability and integration challenges due to limitations in optical access, fabrication techniques, and material compatibility. State-of-the-art quantum computers and atomic clocks are investigating monolithic integration, which necessitates custom traps for each ion species and has not overcome the integration hurdles presented by merging these technologies. The Chimera (Ion Traps and Packaging for Heterogeneous Integration) project proposes a novel approach utilizing heterogeneous integration (HI) of ion traps and photonic circuits. This separation of components allows for flexibility in ion trap design and reduces fabrication compromises. The Chimera project specifically designed an ion trap to interface vertically with a separately fabricated waveguide chip and demonstrates the first steps to integrating them at the packaging level. The ion trap features a large area of removed silicon, allowing the photonics chip outputs closer to the ion trap, improving alignment and packaging processes. The alignment must be accurate to < 1 µm to ensure that the light from the waveguide can overlap with the trapping region. This fine alignment must also be maintained through an ultra-high vacuum bake, a critical step in preparing an ion trap experiment. By combining separate chips, we demonstrate a new path for scaling trapped ion technology that is less reliant on monolithic integration. We successfully fabricated a trap with a large area of oxide removed, resulting in a region thinned to about 40 µm, a key milestone toward successful integration.

42 ENGINEERING↗

Corrosion of 316 L stainless steel under the natural circulation of molten NaCl-MgCl 2 salt

The corrosion behavior of 316 L stainless steel (SS) was studied via the natural circulation of molten eutectic NaCl-MgCl 2 salt through a microloop. The post-corrosion tested 316 L SS microloop sections were characterized with microscopy techniques to determine the microstructural and microchemical changes that occurred at the alloy/salt interface. It was found that 316 L SS showed heterogeneous dissolution at the hot-leg, whereas deposition of corrosion products occurred at the cold-leg. For the first time, experimentally obtained molten salt flow-induced corrosion of 316 L SS results were combined with computational thermodynamic-kinetic models to validate the dissolution and deposition in terms of elemental compositional changes at the alloy/salt interface. The thermodynamic-kinetic modeling predicted that the heterogeneous dissolution of Cr and Fe from the hot-leg section of 316 L SS persisted throughout the salt circulation. The model also estimated that, despite Cr deposition starting earlier than Fe, the total redeposition of Fe is expected to be significantly greater than that of Cr over the circulation of salt. Furthermore, the modeling accurately predicted the subsurface enrichment of Mo which is attributed to the reduced Cr activity and the relatively higher diffusion rate of Mo within the alloy matrix. Here, the agreement between modeling and experimental results confirms that Fe chlorides dissolve at the hot-leg and subsequently deposit at the cold-leg due to activity changes driven by the thermal gradient. By contrast, Cr was not detected in the cold-leg deposits, which is attributed to its weaker temperature dependence on activity, limiting its redeposition under these conditions.

36 - MATERIALS SCIENCE↗

Incorporating Coverage-Dependent Reaction Barriers into First-Principles-Based Microkinetic Models: Approaches and Challenges

Mean-field microkinetic models (MKMs) are appealing for their relatively facile construction, computational tractability, and high-throughput catalyst screening capabilities. As such, they will continue to be a valuable tool for materials design in heterogeneous catalysis even as the field aims to describe more complex systems. Numerous prior reports have provided the groundwork for constructing first-principles-based MKMs, including the analysis of strategies for incorporating lateral interactions into thermodynamic parameters (e.g., adsorption energies). Yet, there remains a need for concerted dialogue on methods for calculating and incorporating coverage-dependent kinetic parameters into MKMs. In this Perspective, we assess strategies for doing so, including the corresponding key physical implications and computational challenges. Here, we emphasize that decoupling thermodynamic and kinetic parameters within MKMs can violate thermodynamic consistency and risk unphysical solutions. For some reactions and catalyst materials, scaling relationships can predict coverage-dependent activation energies, but there are several exceptions evident in the literature, indicating that this approach is not universally applicable and that the field could benefit from research aimed at elucidating the limitations. Conducting high-coverage transition state searches is a rigorous but computationally costly strategy, and the effects of various methods for mitigating this cost on resulting energetics have yet to be broadly explored and validated. The goal of this Perspective is to generate discussion on and inspire focused research into the physical relevance of approaches for describing coverage-dependent reaction barriers in MKMs, including the development of computationally tractable methodologies, to advance the applicability of MKMs across diverse reaction chemistries and conditions.

36 MATERIALS SCIENCE↗

Analytical Models of Frequency and Voltage in Large-Scale All-Inverter Power Systems

Low-order frequency response models for power systems have a decades-long history in optimization and control problems such as unit commitment, economic dispatch, and wide-area control. With a few exceptions, these models are built upon the Newtonian mechanics of synchronous generators, assuming that the frequency dynamics across a system are approximately homogeneous, and assume the dynamics of nodal voltages for most operating conditions are negligible, and thus are not directly computed at all buses. As a result, the use of system frequency models results in the systematic underestimation of frequency minimum nadir and maximum RoCoF, and provides no insight into the reactive power-voltage dynamics. This paper proposes a low-order model of both frequency and voltage response in grid-forming inverter-dominated power systems. The proposed model accounts for spatial-temporal variations in frequency and voltage behavior across a system and as a result, demonstrates the heterogeneity of frequency response in future renewable power systems. Electromagnetic transient (EMT) simulations are used to validate the utility, accuracy, and computational efficiency of these models, setting the basis for them to serve as fast, scalable alternatives to EMT simulation, especially when dealing with very large-scale systems, for both planning and operational studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Coarse Graining Discrete Element Method Information in Particle-in-Cell Length Scales Using a Machine Learning Approach

This report details the development of a machine learning (ML)-driven framework to coarse-grain inter-particle collision dynamics from high-fidelity Discrete Element Method (DEM) simulations to Particle-in-Cell (PIC) scales for gas-solid systems. Traditional PIC models, while computationally efficient, rely on empirical granular stress formulations that fail to capture the full complexity of collision physics, particularly the heterogeneity in particle dynamics. This study adopts a bottom-up approach, integrating insights from DEM simulations to improve the physical fidelity and interpretability of PIC-scale models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A novel ignition model for low velocity impact of heterogeneous explosives based on interacting hot spots

While numerous studies have focused on the ignition of explosives occurring in high velocity impact and the associated shock-to-detonation transition, there has been growing interest in developing computational models focused on low-velocity impact regimes. A predictive low-velocity impact ignition model will be important for analyzing high explosive safety and potential accident scenarios. This work introduces a novel ignition model based on the concept of thermally interacting hot spots to simulate low velocity impacted heterogeneous explosives where observed ignition times are on the order of milliseconds. The model asserts that relevant hot spots are micron-sized, the typical separation between neighboring hot spots is on the order of a hundred microns, and that neighbors interact thermally through heat conduction across the interstitial region between them. To achieve tractable numerical solutions, hot spots are assumed to form a periodic array as opposed to the highly irregular positioning in an actual explosive. This idealization allows a single two hotspot system to characterize the ignition process. Consequently, the model is referred to as the two hot spot Frank-Kamenetskii ignition model. In the present study, hot spots are modeled as constant heat sources terms, but this can be extended to include grain-scale phenomena like frictional heating of micron-sized growing cracks that are confined under high pressure. Because the micron-sized features are below the scale that can be efficiently resolved at a systems level, an efficient subscale scheme based on the Method of Weighted Residuals (MWR) is used to efficiently solve the equations. In conclusion, we carry out numerical examples and analytic predictions illustrating the accuracy and the functioning of the model.

97 MATHEMATICS AND COMPUTING↗

Supporting multiple hardware architectures at CMS: the integration and validation of POWER9

Computing resources in the Worldwide LHC Computing Grid (WLCG) have been based entirely on the x86 architecture for more than two decades. In the near future, however, heterogeneous non-x86 resources, such as ARM, POWER and Risc-V, will become a substantial fraction of the resources that will be provided to the LHC experiments, due to their presence in existing and planned world-class HPC installations. The CMS experiment, one of the four large detectors at the LHC, has started to prepare for this situation, with the CMS software stack (CMSSW) already compiled for multiple architectures. In order to allow for a production use, the tools for workload management and job distribution need to be extended to be able to exploit heterogeneous architectures. Profiting from the opportunity to exploit the first sizable IBM Power9 allocation available on Marconi100 HPC system at CINECA, CMS developed all the needed modifications to the CMS workload management system. After a successful proof of concept, a full physics validation has been performed in order to bring the system in production. The experiences are of very high value, when it comes to commissioning of the similar (even larger) Summit HPC system at Oak Ridge, where CMS is also expecting a resource allocation. Moreover the compute power of those systems is being provided also via GPUs and this represents an extremely valuable opportunity to exploit the offloading capability already implemented in CMSSW. The status of the current integration including the exploitation of the GPUs, the results of the validation as well as the future plans will be shown and discussed.

Boccali, Tommaso [INFN, Pisa]↗

DONUT: physics-aware machine learning for real-time X-ray nanodiffraction analysis

Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Materials science↗

DONUT: Physics-aware Machine Learning for Real-time X-ray Nanodiffraction Analysis

SF-25-088 Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Zhou, Tao [Argonne National Laboratory (ANL), Argo↗

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Atomic-scale identification of active sites of oxygen reduction nanocatalysts

Heterogeneous nanocatalysts play a crucial role in both the chemical and energy industries. Despite substantial advancements in theoretical, computational and experimental studies, identifying their active sites remains a major challenge. Here we utilize atomic electron tomography to determine the three-dimensional atomic structure of PtNi and Mo-doped PtNi nanocatalysts for the electrochemical oxygen reduction reaction. We then employ the experimental atomic structures as input to first-principles-trained machine learning to identify the active sites of the nanocatalysts. Through the analysis of the structure–activity relationships, we formulate an equation termed the local environment descriptor, which balances the strain and ligand effects to provide physical and chemical insights into active sites in the oxygen reduction reaction. The ability to determine the three-dimensional atomic structure and chemical composition of realistic nanoparticles, combined with machine learning, could transform our fundamental understanding of the active sites of catalysts and guide the rational design of optimal nanocatalysts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Celeritas: Accelerating Geant4 with GPUs

Celeritas [1] is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications (specifically, High Lumi- nosity Large Hadron Collider (HL-LHC) simulation) on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a GPU-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.3, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling integration of Celeritas into high energy physics (HEP) experimental frameworks such as CMSSW. On the Summit supercomputer, Celeritas performs EM physics between 6 and 32 faster using the machine’s Nvidia GPUs compared to using only CPUs. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.8–2.3× on GPU and 1.2× on CPU.

Johnson, Seth R.↗

Performance Results on CPU/GPU Exascale Architectures for OMEGA: The Ocean Model for E3SM Global Applications

The US Department of Energy (DOE) conducts climate simulations on some of the world’s largest supercomputers. These exascale machines use heterogeneous architectures with both CPUs and GPUs, and scientific codes must adapt to make full use of this computing power. Los Alamos National Lab is developing Omega: The Ocean Model for E3SM Global Applications, which is specifically designed for modern exascale computers. It uses external libraries that have been optimized for a variety of architectures to run on different supercomputers. Omega is an unstructured-mesh ocean model based on TRiSK numerical methods. It will be the new ocean component of the DOE’s Energy Exascale Earth System Model (E3SM). The algorithms in Omega follow those of the current ocean component, MPAS-Ocean, but it will be written in C++ rather than Fortran to take advantage of the Kokkos performance portability library. Omega spatial operators are written as Kokkos kernels to run efficiently on both CPUs and GPUs. Work on Omega began in 2023 with a new C++ framework for unstructured mesh partitioning, halo exchanges, parallel IO, and Kokkos interfaces. The current version, Omega-0, is being developed to solve the shallow water equations and at present includes all of the tendency terms but not time stepping. Here we share the results of Omega-0 verification and performance testing. Verification includes unit tests implemented with CTest as well as convergence tests in Polaris, an in-house python package with a large suite of test problems. Performance tests compare simulations conducted on CPUs versus GPUs and across different architectures: tests are run on Frontier, which has AMD “Optimized 3rd Gen EPYC” CPUs and AMD MI250X GPUs, as well as Perlmutter, which is composed of AMD EPYC 7763 CPUs and NVIDIA A100 GPUs.

58 GEOSCIENCES↗

The Heterogeneous Integration of Electronic Components

Heterogeneous integration (HI) of electronics components is broadly recognized as a powerful and crucial enabler for the continued growth of computing and communication. From 2010 onwards, the value of HI is increasingly visible in the advanced packaging used in artificial intelligence, high-performance computing, smartphones and communications product implementations. In this Perspective, we argue that HI is crucial to semiconductors and more broadly to the continued evolution of computing and communications. We use leading-edge advanced packaging examples to represent the value, advancements and opportunities for HI. To succeed, it is critical to develop comprehensive HI roadmaps that inform collaborations across the design, manufacturing and reliability spectrum between systems architects, packaging and semiconductor technologists to common goals. Although this article does not provide a full roadmap, we instead detail additional parameters for artificial intelligence, smartphone and other cellular communication devices, and their constituent building blocks including interconnects, power electronics, photonics, thermal management, reliability, modelling and co-design, to foster greater collaboration opportunities among academia, research laboratories and industry.

42 ENGINEERING↗

The impact of capillary heterogeneity on CO 2 flow and trapping across scales

Capillary heterogeneity has been identified over the last decade as a key control on subsurface CO 2 flow behavior during geological CO 2 sequestration. These heterogeneities can be formed in all sedimentary rocks, ranging from slight variations in the sand grain sizes to extensive sequences of interbedded sands, shales, and limestones. Capillary heterogeneity has been largely, although not entirely, overlooked in subsurface flow modeling because it is assumed to only directly influence fluid redistribution over scales of centimeters to meters. However, even small-scale fluid movements can result in dramatic impacts on the mobility and trapping of the CO 2 over kilometers. Therefore, neglecting capillary heterogeneity at multiple scales could potentially lead to errors in modeling and predicting field-scale plume migration. In this review paper, we aim to provide a consistent overview to (1) establish that capillary heterogeneity can have a major impact on CO 2 plume migration, (2) establish the respective length scales at which capillary heterogeneity matters, and (3) provide guidance for numerical modeling. This review covers pertinent literature and extracts key observations from the core to the field scales. Experimental studies have shown that millimeter-decimeter scale capillary heterogeneity can cause the so-called capillary heterogeneity trapping in addition to pore-scale residual trapping. Even at such a small scale, capillary heterogeneity can already lead to complex upscaled constitutive relationships, such as flow-rate dependent and anisotropic relative permeability, which affects field-scale CO 2 migration even when field-scale heterogeneities are present. Under gravity-dominated flow regimes, centimeter-meter scale capillary heterogeneity can entrap a significant amount of CO 2 at field scale, not just after imbibition but also during drainage. In certain cases, the presence of capillary heterogeneity can even completely stop the vertical movement of the CO 2 plume, hence greatly reducing leakage risks. At meter-kilometer scale, the influence of capillary heterogeneity is more pronounced and can hinder or redirect CO 2 migration in both lateral and vertical directions. The impact of capillary heterogeneity across multiple spatial scales poses a great challenge in modeling CO 2 migration at field scale, because it is practically impossible to build a field-scale earth model with grid blocks at millimeter scale. We recommend a hierarchical modeling approach to address this challenge. At field scale, earth models are built to capture geological features and heterogeneities in high but still practical grid resolutions. For each facies or rock type of the field-scale model, high- resolution meter-scale “conceptual” models are built with millimeter-scale grid blocks to capture representative fine-scale bedding geometries and heterogeneities in various environments of deposition, bridging the gap from subcore scale to the size of a field-scale simulation grid block. Upscaling is then used to preserve the smaller-scale flow dynamics of various rock types in field-scale simulations. Here, future work is needed to (1) refine, improve, and validate the hierarchical modeling approach; (2) build libraries of fine-scale bedding models for facies in various environments of deposition; (3) quantify multiscale capillary heterogeneity effects under subsurface uncertainties; (4) gain learning from different storage formations; and (5) establish best practices that balance accuracy and computational speed.

Capillary heterogeneity↗

Physics-informed neural networks for heterogeneous poroelastic media

This study presents a novel physics-informed neural network (PINN) framework for modeling poroelasticity in heterogeneous media with material interfaces. The approach introduces a composite neural network (CoNN) where separate neural networks predict displacement and pressure variables for each material. While sharing identical activation functions, these networks are independently trained for all other parameters. To address challenges posed by heterogeneous material interfaces, the CoNN is integrated with the Interface-PINNs (I-PINNs) framework (Sarma et al., Comput. Methods Appl. Mech. Eng. 429: 117135, 2024), allowing different activation functions across material interfaces. Further, this ensures accurate approximation of discontinuous solution fields and gradients. Performance and accuracy of this combined architecture were evaluated against the conventional PINNs approach, a single neural network (SNN) architecture, and the eXtended PINNs (XPINNs) framework through two one-dimensional benchmark examples with discontinuous material properties. The results show that the proposed CoNN with I-PINNs architecture achieves an RMSE that is two orders of magnitude better than the conventional PINNs approach and is at least 40 times faster than the SNN framework. Compared to XPINNs, the proposed method achieves an RMSE at least one order of magnitude better and is 40% faster.

42 ENGINEERING↗