Search NASASearch

SEARCH · Search NASA

Results for “parallel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field

Using Parameter Sweep in WaterTAP to Analyze New Water Treatment Technologies

We describe a powerful and generalized parameter sweep tool in this report that was originally developed to analyze the performance of existing and novel water treatment models being developed in WaterTAP. Since WaterTAP is built upon IDAES and Pyomo, the parameter sweep tool can be used to systematically explore and debug the behavior of most Pyomo and IDAES numerical models. In order to enable meaningful analyses, the parameter sweep tool has been designed with the following features: 1) Model flexibility: The parameter sweep tool does not enforce any restrictions on the types of models that can be used with it. As long as a Pyomo model can be solved and the parameter is active and mutable, the tool only needs functions that describe how to run the model, the sweep parameters, and the output quantities of interest. 2) Flexible sampling: The parameter sweep tool has inbuilt functions to generate samples from a random distribution or a multidimensional Euclidean space. Furthermore, the users have to ability to supply samples generated from a tool of their choice. 3) Multiple sweep types: A user can choose from one of 3 types of parameter sweeps depending on their needs. 4) Detailed outputs: Outputs generated by the parameter sweep tool can be stored in detailed H5 file or user-friendly CSV files for post processing. 5) Parallel computing: The parameter sweep supports shared and distributed memory parallel computing to enable the use of high performance computers (HPC) for large-scale analyses. 6) Modular: The parameter sweep tool is self-contained and can easily be integrated within an outer-loop analysis or as desired by the user. 7) Ease of use: The tool is well documented and a simple sweep can be easily executed by following the online documentation in a few lines of code. We demonstrate the use of the parameter sweep tool on a simple water treatment system from the WaterTAP repository and show its parallel scaling performance on an Apple laptop and NREL's Eagle HPC. The parameter sweep tool is actively being used with models currently being developed within WaterTAP and we expect its use to grow beyond it to other IDAES and Pyomo models.

97 MATHEMATICS AND COMPUTING

The PARADIGM Project: Case Study in Balancing Experiment Uncertainty with Design simplicity

Accurate nuclear data are required for simulations of many applications including nuclear criticality safety. Actinide nuclear data at intermediate energies (from 1 to 100s of keV) are imprecise and inaccurate, because of scarce differential data, and an insufficient theory approach to capture the structures expected in the data to yield evaluated nuclear data, and lack of integral data for proper validation. This is a known deficiency but has proved challenging to address. More specifically, only 5% of integral experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) benchmark suite address intermediate energies (Fig. 1). Associated calculated effective multiplication factor, k eff , values for these experiments are far outside the experimental uncertainties and are 25× further from experiment than for fast energies. These differences could either stem from systematic biases in nuclear data, experiments or both. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly reduce (by more than tens of percent) the uncertainties of intermediate energy actinide nuclear data. The PARADIGM project designed and intends to execute LANSCE (Los Alamos Neutron Science CEnter) and NCERC (National Criticality Experiments Research Center) intermediate experiments in parallel. They will specifically address a high priority nuclear data need—reducing bias and uncertainty in intermediate plutonium nuclear data. The two experiment will achieve that by informing each other and nuclear theory. By doing all these steps in parallel, the timeline to deliver improved nuclear data to users will significantly be reduced. This work will focus on the integral experiment final design and the balance of design and modeling simplicity while minimizing experiment uncertainty.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Comparative Analysis of Radial and Random Microstructures of Mesophase Pitch Carbon Fibers

Carbon fibers (CF) with radial and random microstructures are produced. Here, these fibers are subjected to identical treatment before being mechanically tested and analyzed with Weibull analysis, with the results revealing a statistically significant difference in tensile strengths of 2.23 GPa for random CF and 1.69 GPa for radial CF. Raman mapping probed the crystalline structure perpendicular to the fiber axis and found a uniform structure, while wide‐angle X‐ray diffraction showed a significant difference of 7.5 Å in the crystallites’ basal lengths parallel to the fiber. Small‐angle X‐ray scattering is completed parallel to the fiber for the first time. A cross‐section Guinier plot of the 1D azimuthal integration is generated assuming symmetric scattering, and the parallel scatterers are found to have a similar length scale to the crystallite's length, validating the testing method. Finally, transmission electron microscopy is completed on the longitudinal cross‐section of each fiber. The radial carbon fiber is found to have a core–shell structure, as evidenced further by fast Fourier transform images. Through all studies, it is shown that the structure developed during mesophase pitch spinning altered the microstructure, thus impacting the mechanical properties, confirming a direct relationship between processing, structure, and properties.

Scherschel, Alexander [Univ. of Virginia, Charlott

Impact of Different Thermal Gradients on the Dynamics of Cylindrical Lithium-ion Cells Subject to Accelerated Aging and on Module Performance

This study investigates the impacts of applying different thermal gradient patterns to cylindrical lithium-ion cells in a module on cell dynamics (temperatures, current flows, state of charge), module performance (evolution of resistance, capacity, and energy versus cycle number), and module lifetime. The thermal gradients were generated using cooling plates (CPs) with three different flow-field designs, namely, straight, perpendicular, and U-turn. The study uses computational fluid dynamics (CFD), the pseudo-two-dimensional (P2D) battery model, capacity loss and increased impedance due to the growth of a solid-electrolyte-interphase, and the electric current distribution from module terminals to cells that depends on the series-parallel electrical connections among the cells. The impact of the thermal gradient (resulting from the CP designs) on the variability in resistance, current, state of charge, and voltage among the cells was analyzed and linked to differences in the module's performance. Applying a thermal gradient to parallel-connected strings of series-connected cells led to variation in the current through each parallel string and an imbalance in the voltage of series-connected cells. Module performance is poorer when the thermal gradient causes a voltage imbalance than when it causes a current imbalance. Module performance becomes the worst when both current variation and voltage imbalance happen together. For instance, the module's lifetime (estimated as reaching 80% of its initial capacity) varied by 5% to 17.5%, depending on the magnitude and pattern of the imposed thermal gradient. As the relative orientation between thermal gradients and cells' electrical connectivity influences the module's performance, appropriate consideration should be given to the choice of the CP, especially if large thermal gradients are allowed.

Battery thermal management

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability

Antielectrophoretic Response-Driven Bending–Tilting Deformation of Cationic Polyelectrolyte Brushes Drives Nonlinear Electroosmotic Transport in Brush-Grafted Nanochannels

In this paper, we use all-atom molecular dynamics (MD) simulations to describe a non-linearly enhanced electroosmotic (EOS) flow, where, in a nanochannel grafted with cationic PMETAC ([Poly(2-(Methacryloyloxy)Ethyl) Trimethylammonium Chloride]) brushes, a two-fold increase in the electric field strength leads to a several-fold (more than two-fold) increase in the EOS flow strength and volume flow rate. The electric field enforces the PMETAC brushes to undergo a bending-tilting driven deformation with a significant portion of the brush layer becoming parallel to the grafting surface. In response, a substantial fraction of the counterions leave the brush layer (hence become more mobile), but instead of going into the bulk, accumulate at the brush-bulk interface, i.e., stay in proximity of the brush segments aligned parallel to the grafting surface. This creates an interesting situation, where the counterions are not completely within the brush layer, yet they fully screen the brush charges. Such “freer” conditions enable the counterions to achieve very high velocity, thereby ensuring that the water solvating the counterions themselves move very fast triggering the significantly augmented EOS transport. Probing deeper we can identify that the bending-tilting driven brush deformation, enforcing the brushes to align parallel to the substrate, results from the anti-electrophoretic behavior of the brushes, where despite being positively charged, the brushes move against the electric field direction. Such an anti-electrophoretic behavior of the PE brushes, which has not been reported before, can be associated with the very fast velocities of the negatively charged counterions and the electrostatic and hydrodynamic coupling of the counterions with the positive functional groups of the brushes. Here, we anticipate that the findings of this paper will shed light on strategies for nanochannel flow, the anti-electrophoretic response of charged polymer chains, and the significance of capturing the detailed chemical architecture of polyelectrolytes in nanoscale science and engineering.

36 MATERIALS SCIENCE

Motion of Molecules in Supramolecular Scaffolds Enhances Bone Regeneration

The regeneration of human tissues is a great scientific challenge and a critical factor to achieve a long healthspan and prevent disabilities due to injury or disease. Materials chemistry can contribute to this goal with the development of bioactive supramolecular systems that can signal cells for regeneration. Recent work in our laboratory using in vivo models of spinal cord injury and cartilage regeneration has demonstrated that the motion of bioactive molecules in supramolecular scaffolds enhances receptor signaling. We report here on a novel molecular strategy to control supramolecular motion in filamentous assemblies using bone regeneration as a functional target. The supramolecular assemblies are composed of monomers that arrange, by design, with either parallel or antiparallel β-sheets, and some of them contain a terminal peptide sequence that binds BMP-2. We found that parallel β-sheet supramolecular assemblies promote greater osteogenic differentiation of progenitor cells in vitro relative to antiparallel assemblies, as well as superior quality of newly regenerated bone in a rat model of spinal fusion. Furthermore, these assemblies drastically reduce the dangerous supraphysiological dose of BMP-2 used clinically for spinal fusion. Here, we attribute the enhanced bioactivity to the weaker nature of hydrogen bonds in parallel relative to antiparallel β-sheet assemblies, which in turn allows greater supramolecular motion and cell signaling of the growth factor-binding molecules.

Anatomy

Massive, long-lived electrostatic potentials in a rotating mirror plasma

Abstract Hot plasma is highly conductive in the direction parallel to a magnetic field. This often means that the electrical potential will be nearly constant along any given field line. When this is the case, the cross-field voltage drops in open-field-line magnetic confinement devices are limited by the tolerances of the solid materials wherever the field lines impinge on the plasma-facing components. To circumvent this voltage limitation, it is proposed to arrange large voltage drops in the interior of a device, but coexist with much smaller drops on the boundaries. To avoid prohibitively large dissipation requires both preventing substantial drift-flow shear within flux surfaces and preventing large parallel electric fields from driving large parallel currents. It is demonstrated here that both requirements can be met simultaneously, which opens up the possibility for magnetized plasma tolerating steady-state voltage drops far larger than what might be tolerated in material media.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

3D nanolithography with metalens arrays and spatially adaptive illumination

The growing demand for advanced materials, miniaturized devices and integrated microsystems calls for the reliable fabrication of complex, multiscale, three-dimensional (3D) architectures, a need increasingly addressed through light-based and laser-based processes. However, owing to the field-of-view (FOV) limitations of conventional imaging optics, existing 3D laser nanofabrication techniques face fundamental challenges in throughput, proximity error and stitching defects on the path to scaling. Here, in this study, we present a scalable 3D nanofabrication platform that uses a metalens-generated focal spot array to parallelize two-photon lithography (TPL) beyond centimetre-scale write field areas. Metalenses are ideally suited for producing submicron-scale focal spots for high-throughput nanolithography, as they uniquely feature large numerical apertures (NAs), immersion media compatibility and large-scale manufacturability. We experimentally demonstrate a printing system that uses a 12-cm 2 metalens array to produce more than 120,000 cooperative focal spots, corresponding to a throughput exceeding 10 8 voxels s −1 . By programmatically patterning the focal spot array using a spatial light modulator (SLM), an adaptive parallel printing strategy is developed for precise greyscale linewidth modulation and choreographed printing of semiperiodic and fully aperiodic 3D geometries. We demonstrate parallel printing of replicated microstructures (>50 M microparticles per day), centimetre-scale 3D architectures with feature sizes down to 113 nm, and photonic and mechanical metamaterials. This work demonstrates the potential of 3D nanolithography towards wafer-scale production, showing how TPL could be used at scale for applications in microelectronics, biomedicine, quantum technology and high-energy laser targets.

Materials science

Sprain energy consequences for damage localization and fracture mechanics

The 2023 smooth Lagrangian Crack-Band Model (slCBM), inspired by the 2020 invention of the gap test, prevented spurious damage localization during fracture growth by introducing the second gradient of the displacement field vector, named the “sprain,” as the localization limiter. The key idea was that, in the finite element implementation, the displacement vector and its gradient should be treated as independent fields with the lowest ( C 0 ) continuity, constrained by a second-order Lagrange multiplier tensor. Coupled with a realistic constitutive law for triaxial softening damage, such as microplane model M7, the known limitations of the classical Crack Band Model were eliminated. Here, we show that the slCBM closely reproduces the size effect revealed by the gap test at various crack-parallel stresses. To describe it, we present an approximate corrective formula, although a strong loading-path dependence limits its applicability. Except for the rare case of zero crack-parallel stresses, the fracture predictions of the line crack models (linear elastic fracture mechanics, phase-field, extended finite element method (XFEM), cohesive crack models) can be as much as 100% in error. We argue that the localization limiter concept must be extended by including the resistance to material rotation gradients. We also show that, without this resistance, the existing strain-gradient damage theories may predict a wrong fracture pattern and have, for Mode II and III fractures, a load capacity error as much as 55%. Finally, we argue that the crack-parallel stress effect must occur in all materials, ranging from concrete to atomistically sharp cracks in crystals.

Science & Technology - Other Topics

1D model of SOL with self consistent calculation of non-coronalradiation and applications to shallow field line angles

A 1D model of the scrape off layer (SOL) is created with a physics based calculation of the impurity radiation and the inclusion of the perpendicular heat flux q ⊥ compared to previous 1D models. The calculation of the impurity radiation includes the addition of a model using parallel and perpendicular motion to calculate the impurity confinement time τ and the inclusion of charge exchange between neutral hydrogen and impurity ions. Reasonable agreement is found between the model and the 2D fluid code SOLPS-ITER. Additionally, an improvement is shown over the constant τ used in previous works The model is used to examine the effects of shallow field line angles at ITER levels of parallel heat flux (q || ). A large increase in radiation results in an increase in the upstream parallel heat flux which can still be in detachment at shallow field line angles for both q ⊥ and the calculated τ for traditional impurities (argon and nitrogen) as well as boron. The physics behind this increase in radiation at shallow field line angles is shown to be a decrease in τ and a change in the temperature profile with the inclusion of q ⊥ .

ITER

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing

Scalable Computation of Topological Abstractions for Scalar Data

Topological data analysis has become an important tool for large scale scalar data analysis and visualization, efficiently extracting the inherent structure and features of interest of the data. However, with growing dataset sizes and complexity, it is increasingly becoming infeasible to compute topological abstractions of interest in serial and on single machines. This paper presents the state of the art in the scalable computation of topological abstractions on scalar data, in shared memory parallel on single machines, and in distributed memory parallel on multiple machines. We highlight results for set‐based, graph‐based and complex‐based abstractions and organize the state of the art based on this taxonomy. The paper identifies parallelization and distribution techniques common in topological algorithms and highlights further areas of interest with underdeveloped efforts.

97 MATHEMATICS AND COMPUTING

Continental rift evolution and drainage reorganization along the Dead Sea rift since the Miocene

The Dead Sea fault is a section of the Arabian-African plate boundary. Widespread field relations indicate that three major drainage systems (stages) occupied the landscape west of the Dead Sea fault since its initiation at ca. 20 Ma. Specifically, (1) an early to middle Miocene drainage system, only minorly reconfigured by the fault (all sediments of this system belong to the Hazeva Formation); (2) a late Miocene to early Pleistocene fault-parallel drainage system named Paran-Neqarot (all sediments of this system belong to the Arava and Zehiha Formations; and (3) the early Pleistocene to present drainage configuration. The temporal and spatial frameworks of drainage stage 1 are generally constrained by radiometric dating of interfingering volcanic units, and the onset and temporal and spatial frameworks of drainage stage 3 are well constrained by cosmogenic 10 Be surface exposure ages. The timing and longevity of the rift-parallel drainage (stage 2) have until now been evasive to direct dating. The overall time gap between stages 1 and 3 is ~12–13 million years. Thus, an early age (within this time gap) of stage 2 would imply an immediate response of drainage reorganization to rift tectonics, while a later age of this drainage system would imply a delayed response. We present 11 10 Be- 26 Al cosmogenic burial ages of alluvial and colluvial units related to the fault-parallel drainage system (stage 2), which collectively constrain the time of deposition of the Arava Formation sediments in the central Negev to ca. 8 Ma. The general lack of stratigraphic order, together with the large dispersion of ages both across and within the sampling sites, attests to significant recycling of sediments from drainage stage 1 into the Arava Formation deposits. The termination of Arava and Zehiha Formation sediment deposition at ca. 1.8 Ma was determined previously using cosmogenic exposure ages of desert pavements that cover the formations. Combining the previously published data with our new data, we established the longevity and character of the Paran-Neqarot drainage system. In conlcusion, this framework highlights the temporal aspect of drainage system build-up and collapse as it responded to transform and extensional plate boundary tectonics during the Neogene.

58 GEOSCIENCES

A Simple, Scalable Large Deformation Solid Mechanics Implementation in the MOOSE Framework

This article describes a large deformation solid mechanics solver implemented as part of the freely available and open source MOOSE finite element simulation framework. The article documents the choices made in developing the solid mechanics framework and describes novel formulations for the gradient operator and constitutive modeling framework made to simplify implementations of different coordinate systems, stabilized gradient operators, and different constitutive model inputs and outputs. In the process, the article describes a new formulation that casts objective integration of the Cauchy stress as a linear transformation of the small stress rate. Finally, the article presents key implementation details and examines the parallel efficiency of the solid mechanics solver implemented in MOOSE. The implementation retains a good weak scaling efficiency beyond 1,000 parallel processes. The article includes a discussion of the factors limiting the parallel efficiency of implicit, large deformation solid mechanics codes on current high-performance computers, with the main current limitation being the scalability of the algebraic multigrid methods used to solve the linearized equilibrium equations.

Applied computing → Computer-aided design

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)