LAMMPS-KOKKOS: Performance portable molecular dynamics across exascale
Presentation at LAMMPS workshop, content identical to recently submitted paper IR 1786384
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Presentation at LAMMPS workshop, content identical to recently submitted paper IR 1786384
Explore the source record for details and available documents.
The quantification of methane concentrations in air is essential for the quantification of methane emissions, which in turn is necessary to determine absolute emissions and the efficacy of emission mitigation strategies. These are essential if countries are to meet climate goals. Large-scale deployment of methane analyzers across millions of emission sites is prohibitively expensive, and lower-cost instrumentation has been recently developed as an alternative. Currently, it is unclear how cheaper instrumentation will affect measurement resolution or accuracy. To test this, the Wireless Autonomous Transportable Methane Emission Reporting System (WATCH 4 ERS) has been developed, comprising four commercially available sensing technologies: metal oxide (MOx,), Non-dispersion Infrared (NDIR), integrated infrared (INIR), and tunable diode laser absorption spectrometer (TDLAS). WATCHERS is the accumulated knowledge of several long-term methane measurement projects at Colorado State University’s Methane Emission Technology Evaluation Center (METEC), and this study describes the integration of these sensors into a single unit and reports initial instrument response to calibration procedures and controlled release experiments. Specifically, this paper aims to describe the development of the WATCH 4 ERS unit, report initial sensor responses, and describe future research goals. Meanwhile, future work will use data gathered by multiple WATCH 4 ERS units to 1. better understand the cost–benefit balance of methane sensors, and 2. identify how decreasing instrumentation costs could increase deployment coverage and therefore inform large-scale methane monitoring strategies. Both calibration and response experiments indicate the INIR has little practical use for measuring methane concentrations less than 500 ppm. The MOx sensor is shown to have a logarithmic response to methane concentration change between background and 600 ppm but it is strongly suggested that passively sampling MOx sensors cannot respond fast enough to report concentrations that change in a sub-minute time frame. The NDIR sensor reported a linear change to methane concentration between background and 600 ppm, although there was a noticeable lag in reporting changing concentration, especially at higher values, and individual peaks could be observed throughout the experiment even when the plumes were released 5 s apart. The TDLAS sensor reported all changes in concentration but remains prohibitively expensive. Our findings suggest that each sensor technology could be optimized by either operational design or deployment location to quantify methane emissions. The WATCH4ERS units will be deployed in real-world environments to investigate the utility of each in the future.
We introduce an extension to the AthenaK code for general-relativistic magnetohydrodynamics (GRMHD) in dynamical spacetimes using a 3+1 conservative Eulerian formulation. Like the fixed-spacetime GRMHD solver, we use standard finite-volume methods to evolve the fluid and a constrained-transport scheme to preserve the divergence-free constraint for the magnetic field. We also utilize a first-order flux correction (FOFC) scheme to reduce the need for an artificial atmosphere and optionally enforce a maximum principle to improve robustness. We demonstrate the accuracy of AthenaK using a set of standard tests in flat and curved spacetimes. Using a SANE accretion disk around a Kerr black hole, we compare the new solver to the existing solver for stationary spacetimes using the so-called "HARM-like" formulation. We find that both formulations converge to similar results. We also include the first published binary neutron star (BNS) mergers performed on graphical processing units (GPUs). Thanks to the FOFC scheme, our BNS mergers maintain a relative error of $\mathcal{O}$(10 –11 ) or better in baryon mass conservation up to collapse. Finally, we perform scaling tests of AthenaK on OLCF Frontier, where we show excellent weak scaling of ≥80% efficiency up to 32,768 GPUs and 74% up to 65,536 GPUs for a GRMHD problem in dynamical spacetimes with six levels of mesh refinement. AthenaK achieves an order-of-magnitude speedup using GPUs compared to CPUs, demonstrating that it is suitable for performing numerical relativity problems on modern exascale resources.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Camera-ready paper and appendix for the Vivek Sarkar Festschrift Symposium.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
U.S. Forest Service (USFS) wildfire base camps use portable power for electrical needs, including yurts and trailers from which logistics staff work during the incident. These yurts and trailers are conventionally powered by portable diesel generators. Hybrid portable power systems consisting of a combination of solar photovoltaics (PV), battery energy storage systems (BESS), and/or backup diesel generators have been used in recent years and were piloted by the National Technology and Development Program (NTDP) on incidents in fall 2024 as part of NTDP's Portable Power Project. As part of that project, this work included the development of two Excel-based tools and two reports explaining the tools. The first tool, the Incident Energy Systems Model, is an Excel-based tool that models the power output of hybrid portable power systems powering yurts and/or trailers at fire camps over a typical day. The associated report (published separately) outlines a user guide for the model and walks through two scenarios. The two scenarios are 1) trailers powered by rooftop solar PV, batteries, and a back-up diesel generator, and 2) trailers powered by ground mount solar PV, batteries, and a back-up diesel generator. The second tool is a high-level cost-benefit analysis of the same two types of portable power systems, and is location independent.
U.S. Forest Service (USFS) wildfire base camps use portable power for electrical needs, including yurts and trailers from which logistics staff work during the incident. These yurts and trailers are conventionally powered by portable diesel generators. Hybrid portable power systems consisting of a combination of solar photovoltaics (PV), battery energy storage systems (BESS), and/or backup diesel generators have been used in recent years and were piloted by the National Technology and Development Program (NTDP) on incidents in fall 2024 as part of NTDP's Portable Power Project. As part of that project, this work included the development of two Excel-based tools and two reports explaining the tools. The first tool, the Incident Energy Systems Model, is an Excel-based tool that models the power output of hybrid portable power systems powering yurts and/or trailers at fire camps over a typical day.
A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.
CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.
There is growing recognition that Demand Flexibility (DF) can play a major role in enhancing grid reliability, with building control applications emerging as key enablers for DF. However, the traditional approach to deploying new control applications in buildings, including those for DF, remains largely manual and tailored to individual buildings, making it difficult to scale. While research efforts have explored semantics-driven portability, DF controls specification, and assessment approaches, these initiatives are fragmented and limited in scope. This paper proposes a novel methodology, grounded in design science research, to integrate these elements and create a comprehensive DF controls library for both industry and academia. This approach is applied to develop the Demand FLEXibility controls LIBrary using Semantics (DFLEXLIBS), an extensible open-source library that provides DF controls for HVAC systems in Python. DFLEXLIBS enables portable, easy-to-deploy controls that abstract building-specific data points, facilitating assessment across diverse buildings. DFLEXLIBS features nine different control applications, and it is successfully implemented and tested across four virtual and two real buildings, bridging the gap between semantics-driven portability, DF controls specification, and rigorous performance assessment. Its benefits are measured by a reusability ratio greater than 90% and a functional overlap ratio of around 70% for the most common functions used in the library, significantly reducing time for deploying new controls.
The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.
This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.
We present the numerical methods, programming methodology, verification, and performance assessment of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures, in this work. Our plasma fluid model solves the coupled conservation equations for species transport, electrostatic Poisson and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive-grid/particle management library, AMReX, and is portable over widely available vendor specific GPU architectures. We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth-order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on capacitive discharges and atmospheric pressure streamer propagation. We demonstrate the use of our solver on two 3D simulation cases: an atmospheric streamer propagation in Ar-H2 mixtures and a low pressure three-electrode radio frequency reactor. Our performance studies on three different CPU+GPU architectures indicate ~ 150-400X speed-up using AMD and NVIDIA GPUs per time step compared to a single CPU core for a 4 million cell simulation with 15 species.