Search NASA⌕ Search

SEARCH · Search NASA

Results for “grain architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Generating multi-scale Li-ion battery cathode particles with radial grain architectures using stereological generative adversarial networks

Abstract Understanding structure-property relationships of Li-ion battery cathodes is crucial for optimizing rate-performance and cycle-life resilience. However, correlating the morphology of cathode particles, such as in LiNi0.8Mn0.1Co0.1O2 (NMC811), and their inner grain architecture with electrode performance is challenging, particularly, due to the significant length-scale difference between grain and particle sizes. Experimentally, it is not feasible to image such a high number of particles with full granular detail. A second challenge is that sufficiently high-resolution 3D imaging techniques remain expensive and are sparsely available at research institutions. Here, we present a stereological generative adversarial network-based model fitting approach to tackle this, that generates representative 3D information from 2D data, enabling characterization of materials in 3D using cost-effective 2D data. Once calibrated, this multi-scale model can rapidly generate virtual cathode particles that are statistically similar to experimental data, and thus is suitable for virtual characterization and materials testing through numerical simulations. A large dataset of simulated particles with inner grain architecture has been made publicly available.

25 ENERGY STORAGE↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

Unconventional compute methods and future challenges for superconducting digital computing

Superconducting digital computing (SDC) based on Josephson junctions (JJs) offers significant potential for enhancing compute throughput and reducing energy consumption compared to conventional room-temperature CMOS-based approaches. Current superconducting logic families exhibit diverse characteristics in clocking strategies, power management, and information encoding techniques. This paper reviews recent advancements in unconventional computing methods specifically designed for superconducting digital circuits, emphasizing temporal computing and pulse-train representations. Notable techniques include race logic (RL), temporal pulse train computing (U-SFQ), and temporal multipliers, each offering unique performance and area advantages suited to superconducting implementations. Additionally, this paper reviews innovations in superconducting coarse-grain reconfigurable architectures (CGRA), superconducting-specific on-chip communication architectures, cryogenic sensor interfaces, and quantum computing control electronics. Finally, we highlight research challenges that should be addressed to facilitate the widespread adoption of superconducting digital computing.

EDA tools↗

Evolution of connectivity architecture in the Drosophila mushroom body

Brain evolution has primarily been studied at the macroscopic level by comparing the relative size of homologous brain centers between species. How neuronal circuits change at the cellular level over evolutionary time remains largely unanswered. Here, using a phylogenetically informed framework, we compare the olfactory circuits of three closely related Drosophila species that differ in their chemical ecology: the generalists Drosophila melanogaster and Drosophila simulans and Drosophila sechellia that specializes on ripe noni fruit. We examine a central part of the olfactory circuit that, to our knowledge, has not been investigated in these species—the connections between projection neurons and the Kenyon cells of the mushroom body—and identify species-specific connectivity patterns. We found that neurons encoding food odors connect more frequently with Kenyon cells, giving rise to species-specific biases in connectivity. These species-specific connectivity differences reflect two distinct neuronal phenotypes: in the number of projection neurons or in the number of presynaptic boutons formed by individual projection neurons. Finally, behavioral analyses suggest that such increased connectivity enhances learning performance in an associative task. Our study shows how fine-grained aspects of connectivity architecture in an associative brain center can change during evolution to reflect the chemical ecology of a species.

59 BASIC BIOLOGICAL SCIENCES↗

Achieving strength-ductility synergy in hierarchical aluminum metal matrix composites via friction extrusion

We report the fabrication of aluminum metal matrix composites (Al-MMCs) with hierarchical architectures via friction extrusion (FE), a scalable, single-step, solid-phase processing technique. Precursor pucks containing 0–15 vol% Al₂O₃ particles were extruded into fully dense AA6061-based composite rods. The FE induced a tree-ring-like architecture of concentric particle-rich and particle-lean bands, yielding refined grains in particle-rich regions and coarser grains elsewhere. At the nanoscale, magnesium in AA6061 selectively reacted with Al₂O₃ particles to form virus-like nodes, improving particle–matrix bonding. This multi-scale design strategy, combining mesoscale architecture, microscale grain refinement, and nanoscale interface engineering overcome the conventional strength–ductility trade-off. Tensile testing showed substantial increases in yield and ultimate tensile strengths while retaining high ductility ( > 20%). Enhanced strain hardening, driven by the accumulation of geometrically necessary dislocations at interfaces, contributed to the performance. The hierarchical microstructure produced by FE demonstrates a promising pathway for scalable fabrication of lightweight MMCs for structural applications requiring a combined high strength and ductility.

Kalsar, Rajib [Pacific Northwest National Laborato↗

Topology-Informed Design Rules for Deconstructable Thermoset Copolymer Networks

Existing models of thermoset deconstruction facilitated by incorporating cleavable comonomers rely on a mean-field reverse gel point paradigm, which predicts network dissolution once cleavable bonds reach a critical stoichiometric threshold, but does not account for where those bonds reside within the network architecture. Using reactive coarse-grained molecular dynamics simulations coupled with graph-theoretic analysis, we extend this stoichiometric picture to show that deconstructability is governed by the curing-imprinted network topology rather than stoichiometry alone. This topological organization is hierarchical: at the local scale, the elastic effectiveness of cross-link junctions determines which cross-links constitute the load-bearing scaffold; at the mesoscale, the cross-linking rate kinetically templates that scaffold into topologically modular communities─densely cross-linked clusters connected by sparse bridging strands that sustain network connectivity. Using betweenness centrality to identify nodes that disproportionately lie on intercommunity shortest paths, we demonstrate that effective deconstruction of the network into macromolecular fragments requires cleavable comonomers to intercept these high-centrality bridging strands. We further find that under uniform, disassortative comonomer incorporation, this topological requirement provides a mechanistic basis for extending the reverse gel point to incorporate network topology. We also show that modularity imposes a fundamental limit on fragment uniformity that persists even when the centrality requirement is met. Finally, we demonstrate that chain stiffness provides a nearly independent lever to suppress mechanically redundant cross-links and raise the glass transition temperature without significantly altering the deconstruction outcome. Together, these findings reframe the thermoset design space around network topology and provide actionable guidelines for engineering thermoset copolymers with predictable deconstructability and targeted thermomechanical performance.

coarse-grained molecular dynamics↗

New Deformation Mechanisms in Nanocrystalline Nano-porous Small Scale Metals as Defined by Kinetically-driven Microstructures

This report describes recent advances in the design and fabrication of nanostructured metallic pillars with hierarchical microstructures using a novel nanoscale additive manufacturing approach. Through a hydrogel-infusion-based two-photon lithography (TPL) method, we successfully fabricated 3D nickel nanopillars exhibiting both nanocrystalline and nanoporous features. The resulting “bamboo-like” internal architecture comprises 30–50 nm grains and voids with similarly scaled pores. These geometrically tunable pillars, with diameters ranging from ~130 to 550 nm, serve as an ideal platform for probing deformation mechanisms in structurally heterogeneous metals at the nanoscale.

36 MATERIALS SCIENCE↗

Secure API-Driven Research Automation to Accelerate Scientific Discovery

The Secure Scientific Service Mesh (S3M) provides API-driven infrastructure to accelerate scientific discovery through automated research workflows. By integrating near real-time streaming capabilities, intelligent workflow orchestration, and fine-grained authorization within a service mesh architecture, S3M enables secure and flexible programmatic access to high performance computing (HPC) resources. This framework allows intelligent agents and experimental facilities to dynamically provision resources and execute complex workflows, accelerating experimental lifecycles, and enabling AI-augmented autonomous science. S3M establishes a modern foundation for scientific computing infrastructure that significantly reduces traditional barriers between researchers, computational resources, and experimental facilities.

Skluzacek, Tyler [ORNL] (ORCID:0000000322424931)↗

Understanding the Influence of Chain Architecture on the Transport Quantities of Polymer Electrolytes with Covalently Bonded Anions

Here, we use a combination of experiments and coarse-grained molecular dynamics simulations to elucidate the structure–property relationships in polymer electrolytes obtained by the copolymerization of poly(vinyl ethylene carbonate─lithium styrene bis(trifluoromethanesulfonyl)imide) or p(VEC-LiSTFSI). Experiments show that the conductivity reduces with increasing anion (i.e., STFSI) fraction on the chain, and the cation transference number (t + ) is found to be dependent on the anion fraction. Furthermore, a significant fraction of unpolymerized VEC monomers are observed. Since it is inherently difficult to experimentally control the chain architecture and the amount of unpolymerized VEC in these systems, we perform coarse-grained molecular dynamics simulations on model polymer systems with different chain architectures to mimic the plausible experimental systems. Specifically, we look at the differences in transference numbers arising from (i) a random copolymer of VEC and STFSI monomers; (ii) a blend of VEC-STFSI copolymer with VEC monomers; and (iii) a ternary blend of the VEC homopolymer, STFSI homopolymers, and VEC monomers. The ternary blend model demonstrates the closest resemblance with the experimental transference numbers and diffusivities. The lithium diffusivity obtained from the coarse-grained models with VEC monomers (plasticizers) is about 1.5 times that of the model without VEC monomers, showing that the plasticizing effect of VEC monomers is modest. We rationalize the experimental observations based on aggregate and cluster analyses obtained from molecular simulations. This work reveals that polymer electrolyte chain architecture and plasticizers can critically influence the transport properties, and these parameters should be considered when designing single ion conducting polymeric electrolytes.

cluster distribution↗

Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments

In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.

59 BASIC BIOLOGICAL SCIENCES↗

Nano-Engineered Interfaces in Dual-Layer Electrodes for Protonic Ceramic Cells with Enhanced Stability and Kinetics

Enhancing interfacial stability and charge transfer in protonic ceramic cells (PCCs) remains a critical challenge, as structural degradation and interfacial resistance often compromise durability and efficiency. Here, we report a nanoengineered dual-layer oxygen electrode architecture designed to address these limitations by introducing a fine-grained nanoparticle interfacial contact layer beneath a porous catalytic backbone. The nanoscale powders, through enhanced sintering activity, densify into a robust interfacial layer that promotes strong chemical bonding, uniform adhesion, and continuous ionic/electronic pathways with the BCZYYb electrolyte. This hierarchical architecture mitigates delamination, redistributes mechanical stress, and establishes efficient charge and mass transport channels without relying on corrosive surface treatments. Electrochemical evaluation demonstrates that the dual-layer design markedly reduces interfacial polarization resistance and accelerates electrode kinetics. Compared to the single-layer counterpart, the architecture achieves a peel strength of 44.53 N/cm 2 , a 40% improvement in peak power density (0.96 W cm –2 at 600 °C), and a 130% enhancement in electrolysis current density (4.78 A cm –2 at 1.57 V). Faradaic efficiency remains as high as 88% under high steam concentrations, underscoring minimal charge loss during practical operation. Notably, the electrode retains stability across 450–600 °C and under transient voltage cycling, with impedance spectra confirming suppressed interfacial resistance growth over prolonged use. These results highlight nanoscale interface engineering as a powerful route to enhance both mechanical robustness and electrochemical kinetics in PCCs. The demonstrated scalability and durability of this architecture provide a versatile platform for advancing solid-state electrochemical systems, including reversible fuel cells and high-efficiency hydrogen production technologies.

Faradaic efficiency↗

An end-to-end workflow for executing a classically bootstrapped variational quantum algorithm on an academic quantum computer

Academic quantum computing platforms often face unique challenges in executing quantum workloads due to fragmented software environments and limited engineering support. Unlike commercial ecosystems, academic devices typically evolve without full-stack integration in mind, making it difficult to run complex applications—such as variational quantum algorithms (VQA)—reliably and efficiently. Issues such as incompatible software layers and lack of automated job management significantly increase the overhead of theory-experiment collaboration. To address these challenges, we develop a modular, end-to-end workflow that decouples application-layer code from low-level hardware control, automates circuit submission and result collection, and supports fine-grained circuit-level job scheduling and recovery. The architecture employs a dual-end application programming interface (API) design, enabling robust operation across unstable or resource-constrained hardware backends. For practical use, the framework is lightweight and user-friendly, allowing rapid prototyping of full-stack workflows using basic Python tools. We validate this workflow on a high-fidelity trapped-ion quantum computer by demonstrating a variational quantum eigensolver (VQE) experiment with a classically bootstrapped ansatz initialization technique. The system successfully executed over 60,000 circuits across multiple molecular test cases with minimal human intervention, highlighting the framework’s effectiveness in enabling reproducible, resilient quantum experimentation in academic settings.

Clifford↗

On a Simplified Approach to Achieve Parallel Performance and Portability Across CPU and GPU Architectures

This paper presents software advances to easily exploit computer architectures consisting of a multi-core CPU and CPU+GPU to accelerate diverse types of high-performance computing (HPC) applications using a single code implementation. The paper describes and demonstrates the performance of the open-source C++ matrix and array (MATAR) library that uniquely offers: (1) a straightforward syntax for programming productivity, (2) usable data structures for data-oriented programming (DOP) for performance, and (3) a simple interface to the open-source C++ Kokkos library for portability and memory management across CPUs and GPUs. The portability across architectures with a single code implementation is achieved by automatically switching between diverse fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. The MATAR library solves many longstanding challenges associated with easily writing software that can run in parallel on any computer architecture. This work benefits projects seeking to write new C++ codes while also addressing the challenges of quickly making existing Fortran codes performant and portable over modern computer architectures with minimal syntactical changes from Fortran to C++. We demonstrate the feasibility of readily writing new C++ codes and modernizing existing codes with MATAR to be performant, parallel, and portable across diverse computer architectures.

97 MATHEMATICS AND COMPUTING↗

Laser Powder Directed Energy Deposition of Steels for Nuclear Applications

This comprehensive investigation examines the structure–property relationships in two nuclear alloy systems—Alloy 709 (A709) austenitic stainless steel and Grade 92 (G-92) ferritic/martensitic (F/M) steel—manufactured via directed energy deposition (DED) for sodium-cooled fast reactor applications. This study establishes the fundamental mechanisms for controlling microstructures for optimizing the mechanical performance of additively manufactured nuclear materials through systematic heat treatment optimization and multiscale characterization. As-deposited A709 steel develops a complex multiscale strengthening architecture consisting of a fine cellular solidification structure with diameter of 2-3 µm within10–50 µm grains, elevated dislocation densities from rapid thermal cycling, and grain boundary precipitates that activate concurrent Hall–Petch, dislocation, and precipitation hardening mechanisms to achieve exceptional properties [yield strength (YS): 603 MPa, ultimate tensile strength (UTS): 844 MPa, Vickers hardness: 220 HV] that achieve a 44% superior strength compared to that of the wrought material. Heat treatments produce different results. Solution annealing (SA) dissolves the cellular structure and reduces the hardness to 190 HV. Precipitation treatment (PT) keeps the cellular structure but adds carbides, allowing the hardness to reach 205 HV. The best approach combines both treatments (SA+PT) and creates uniform precipitate distributions with M 23 C 6 carbides at the grain boundaries and MX carbonitrides in the matrix, achieving a hardness of 195 HV. However, directional differences persist, with a 12%–15% strength variation between orientations due to the inherited layered microstructural architecture that survives aggressive heat treatment. While tensile testing at 550°C demonstrates 40%–50% thermal softening with dynamic strain aging, DED A709 steel still maintains a 71% higher YS than that of the wrought material. Ion irradiation studies (100–400 dpa) of DED A709 steel reveal progressive radiation damage with increasing void density and radiation-induced segregation causing nickel enrichment and chromium depletion, which will ultimately compromise mechanical properties. As-deposited G-92 exhibits exceptional strength (UTS: 1650–1700 MPa, 430 HV) through a complex microstructure containing both ferrite and martensite phases, a high geometrically necessary dislocation (GND) density (17.04×10 14 /m 2 ), and fine carbides. Heat treatments create distinct changes. Normalizing produces fresh martensite with the highest hardness (460 HV) and an increased GND density (20.23×10 14 /m 2 ). Tempering develops dual precipitation systems and reduces the hardness to 290 HV. The optimal approach uses sequential normalizing plus tempering, achieving balanced properties with the lowest hardness (250 HV) and a reduced GND density (11.01×10 14 /m 2 ). A processing-dependent anisotropy is observed: horizontal specimens achieve superior ductile behavior, while vertical specimens exhibit brittle failure. A tempering heat treatment successfully mitigates this anisotropic behavior by transforming the hard martensitic as-deposited structure into tempered martensite enabling both horizontal and vertical specimens to exhibit similar stress–strain characteristics with visible necking behavior. Remarkably, testing at 550°C reveals a reversal in the anisotropy, where as-deposited specimens achieve near isotropy with superior thermal stability (a 15%–20% strength reduction), while tempered specimens develop an orientation dependence with a 25%–30% strength reduction. Both alloy systems demonstrate that DED processing creates specimens with a superior strength through refined microstructural features, though with distinct strengthening mechanisms—austenitic through cellular structures and precipitates versus F/M through phase transformations and precipitates. Heat treatment optimization requires alloy-specific approaches, with A709 benefiting from controlled precipitation while G-92 requires careful phase transformation control. The results show that DED manufacturing can produce nuclear materials with exceptional performance, but directional effects and temperature-dependent behavior must be carefully considered for reactor component design and qualification.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluating Function-as-a-Service (FaaS) frameworks for the Accelerator Control System

As particle accelerator control systems evolve in complexity and scale, the need for responsive, scalable, and cost-effective computational infrastructure becomes increasingly critical. Function-as-a-Service (FaaS) offers an alternative to traditional monolithic architecture by enabling event-driven execution, automatic scaling, and fine-grained resource utilization. This paper explores the applicability and performance of FaaS frameworks in the context of a modern particle accelerator control system, with the objective of evaluating their suitability for short lived and triggered workloads. In this paper, we evaluate prominent open-source FaaS platforms in executing functional logic, triggers, and diagnostics routines. Evaluation metrics consist of cold-start latency, scalability, performance, integration with other open-source tools like Kafka. Experimental workloads were designed to simulate real-world control tasks when implemented as stateless FaaS functions. These workloads were benchmarked under various invocation loads and network conditions. Self-hosted FaaS platforms, when deployed within accelerator networks, offer greater control over execution environment, better integration with legacy systems, and support for real-time guarantees when paired with message queues. Based on lessons learned and evaluation metrics, this paper describes reliability of the FaaS framework for the Accelerator Control Systems (ACS).

Jaikar, A. [Fermilab] (ORCID:0000000332046217)↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Key Factors in Semi-Generic Coarse-Grained Modeling of Solid Polymer Electrolytes

Given the long length and time scales of interest and strong interactions present in solid polymer electrolytes, coarse-grained molecular models can be useful in understanding the molecular basis of their structure and transport properties. Rather than representing atoms individually, coarse-grained models use beads representing groups of atoms, making the system simpler and more efficient. Generic coarse-grained models, that are not built based on matching a particular atomistic system, can provide insight into the important physical considerations that apply across different chemical systems, though it can be unclear how to best match to experimental systems and capture relevant experimental behaviors with as simple of a model as possible. Here, we build on generic bead-spring models, but use stiff angle potentials, set different bead properties for different polymer types, and include additional ion parameters, creating a semi-generic coarse-grained model that can map more specifically to polymers and copolymers with different chain architectures, component glass transition temperatures, and ion solvation behaviors. We set most parameters with basic homopolymer data such as glass transition temperatures, Kuhn length, density, and dielectric constant, with further adjustment based on data from the polymer electrolyte system. We specifically model polystyrene-block-poly(oligo-oxyethylene methyl ether methacrylate) (PS-b-POEM) with lithium triflate salt, considering the POEM to be made of a poly(methylmethacrylate) backbone with poly(ethylene oxide) side chains. Solvation of ions is accounted for by additional polymer-ion interactions of the form -S/r4, plus additional lithium-polymer Lennard-Jones interactions. We find these potentials have different effects and discuss strategies for setting these parameters.

Zhang, Yuanhao↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗