Search NASA⌕ Search

SEARCH · Search NASA

Results for “MFEM”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

High-performance finite elements with MFEM

The MFEM (Modular Finite Element Methods) library is a high-performance C++ library for finite element discretizations. MFEM supports numerous types of finite element methods and is the discretization engine powering many computational physics and engineering applications across a number of domains. Furthermore, this paper describes some of the recent research and development in MFEM, focusing on performance portability across leadership-class supercomputing facilities, including exascale supercomputers, as well as new capabilities and functionality, enabling a wider range of applications. Much of this work was undertaken as part of the Department of Energy’s Exascale Computing Project (ECP) in collaboration with the Center for Efficient Exascale Discretizations (CEED).

97 MATHEMATICS AND COMPUTING↗

MAPS: the MFEM Anisotropic Plasma Solver

Simulating magnetically confined fusion plasmas presents a uniquely challenging problem due to the nonlinear anisotropic heat conduction. We introduce the MAPS (MFEM Anisotropic Plasma Solver) tool, which uses a high-order finite element method to compute transport solutions on unstructured meshes. We show results for a set of three 2-D verification tests, two of which demonstrate the expected convergence properties for various mesh resolutions and polynomial degrees. We then discuss the convergence rate for the third test.

Barnett, Rhea [ORNL] (ORCID:0000000317527979)↗

Finite elements for Matérn-type random fields: Uncertainty in computational mechanics and design optimization

This work highlights an approach for incorporating realistic uncertainties into scientific computing workflows based on finite elements, focusing on prevalent applications in computational mechanics and design optimization. We leverage Matérn-type Gaussian random fields (GRFs) generated using the SPDE method to model aleatoric uncertainties, including environmental influences, variating material properties, and geometric ambiguities. Our focus lies on delivering practical GRF realizations that accurately capture imperfections and variations and understanding how they impact the predictions of computational models as well as the shape and topology of optimized designs. Here we describe a numerical algorithm based on solving a generalized SPDE to sample GRFs on arbitrary meshed domains. The algorithm leverages established techniques and integrates seamlessly with the open-source finite element library MFEM and associated scientific computing workflows, like those found in industrial and national laboratory settings. Our solver scales efficiently for large-scale problems and supports various domain types, including surfaces and embedded manifolds. We showcase its versatility through biomechanics and topology optimization applications, emphasizing the potential to influence these domains. The flexibility and efficiency of SPDE-based GRF generation empowers us to run large-scale optimization problems on 2D and 3D domains, including finding optimized designs on embedded surfaces, and to generate design features and topologies beyond the reach of conventional techniques. Moreover, these capabilities allow us to model and quantify geometric uncertainties on reconstructed submanifolds, such as the interpolated surfaces of cerebral aneurysms provided by postprocessing CT scans. In addition to offering benefits in these specific domains, the proposed techniques transcend specific applications and generalize to arbitrary forward and backward problems in uncertainty quantification involving finite elements.

97 MATHEMATICS AND COMPUTING↗

Benchmark between antenna code TOPICA, RAPLICASOL and Petra-M for the ICRH ITER antenna

ITER will be equipped with three plasma heating systems: neutral beam (NB), electron cyclotron (EC), and ion cy-clotron resonance heating (ICRH). The latter consists of two identical ICRH antennas to deliver 20 MW to the plasma (baseline, upgradable to 40 MW). ICRH will play a crucial role in the ignition and sustainment of burning plasmas in ITER. A high fidelity and robust modeling effort to understand the interaction of the IC waves with the scrape-off-layer (SOL) plasma is a very important aspect. Among the main important research topics, we have the assessment of the antenna loading for different plasma scenarios, the role of the lower hybrid resonance in front of the antenna and how to include it in our models, and the RF sheath boundary conditions to evaluate the antenna impurity generation. In this work, we tackle the first of these by reporting on ICRF simulations employing the Petra-M code, which is an electromagnetic simulation tool for modeling RF wave propagation based on MFEM [http://mfem.org] for the ITER ICRH antenna. Moreover, a benchmark between the well tested antenna codes TOPICA, RAPLI-CASOL, which is based on COMSOL [www.comsol.com], and the Petra-M code is also presented. In conclusion, S- and Z-matrices and wave electric field are compared showing an excellent agreement among these codes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Using the Stix finite element RF code to investigate operation optimization of the ICRF antenna on Alcator C-Mod

Abstract As the Ion Cyclotron Radio Frequency range (ICRF) heating becomes more favorable in fusion devices, the urgency of predicting and mitigating impurity generation that arises from it becomes more pressing. In the ICRF regime, rectified Radio Frequency (RF) sheaths are known to form at antenna and material edges that influence negative effects like sputtering and a decrease in heating efficiency. Methods to mitigate the formation of these RF sheaths through RF image currents cancellation have been experimentally studied. A power-phasing scan done on Alcator C-Mod in which the amount of power on the two inner straps ( P in ) versus the total 4 straps ( P tot ) was varied showed a minimization of enhanced potentials between P in / P tot ∼ 0.7–0.9 while impurities were minimized for P in / P tot ∼ 0.5–0.8. New capabilities in the realm of representing the RF sheath numerically now allow for these experiments to be simulated. Given the size of the sheath relative to the scale of the device, it can be approximated as a Boundary Condition (BC). A new parallelized cold-plasma wave equation solver called Stix implements a non-linear sheath impedance model BC formulated by Myra et al (2015 Phys. Plasmas 22 062507) through the method of finite elements using the MFEM library [ http://mfem.org ]. It is seen that Stix shows qualitative agreement with the measured C-Mod enhanced potentials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance Portable Graphics Processing Unit Acceleration of a High-Order Finite Element Multiphysics Application

The Lawrence Livermore National Laboratory (LLNL) will soon have in place the El Capitan exascale supercomputer, based on advanced micro devices (AMD) graphics processing units (GPUs). As part of a multiyear effort under the National Nuclear Security Administration (NNSA) Advanced Simulation and Computing (ASC) program, we have been developing marbl, a next generation, performance portable multiphysics application based on high-order finite elements. In previous years, we successfully ported the Arbitrary Lagrangian–Eulerian (ALE), multimaterial, compressible flow capabilities of marbl to nvidia GPUs as described in Vargas et al. Here, in this paper, we describe our ongoing effort in extending marbl's GPU capabilities with additional physics, including multigroup radiation diffusion and thermonuclear burn for high energy density physics (HEDP) and fusion modeling. We also describe how our portability abstraction approach based on the raja Portability Suite and the mfem finite element discretization library has enabled us to achieve high performance on AMD based GPUs with minimal effort in hardware-specific porting. Throughout this work, we highlight numerical and algorithmic developments that were required to achieve GPU performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Low-Order Preconditioning for the High-Order Finite Element de Rham Complex

Here, we present a unified framework for constructing spectrally equivalent low-order-refined discretizations for the high-order finite element de Rham complex. This theory covers diffusion problems in H 1 , H(curl), and H(div) and is based on combining a low-order discretization posed on a refined mesh with a high-order basis for Nédélec and Raviart–Thomas elements that makes use of the concept of polynomial histopolation (polynomial fitting using prescribed mean values over certain regions). This spectral equivalence, coupled with algebraic multigrid methods constructed using the low-order discretization, results in highly scalable matrix-free preconditioners for high-order finite element problems in the full de Rham complex. Additionally, a new lowest-order (piecewise constant) preconditioner is developed for high-order interior penalty discontinuous Galerkin (DG) discretizations, for which spectral equivalence results and convergence proofs for algebraic multigrid methods are provided. In all cases, the spectral equivalence results are independent of polynomial degree and mesh size; for DG methods, they are also independent of the penalty parameter. These new solvers are flexible and easy to use; any “black-box” preconditioner for low-order problems can be used to create an effective and efficient preconditioner for the corresponding high-order problem. A number of numerical experiments are presented, based on an implementation in the finite element library MFEM. A range of challenging three-dimensional problems are used to corroborate the theoretical properties and demonstrate the flexibility and scalability of the method.

97 MATHEMATICS AND COMPUTING↗

Real-Time Bayesian Inference at Extreme Scale: A Digital Twin for Tsunami Early Warning Applied to the Cascadia Subduction Zone

We present a Bayesian inversion-based digital twin that employs acoustic pressure data from seafloor sensors, along with 3D coupled acoustic–gravity wave equations, to infer earthquake-induced spatiotemporal seafloor motion in real time and forecast tsunami propagation toward coastlines for early warning with quantified uncertainties. Our target is the Cascadia subduction zone, with one billion parameters. Computing the posterior mean alone would require 50 years on a 512 GPU machine. Instead, exploiting the shift invariance of the parameter-to-observable map and devising novel parallel algorithms, we induce a fast offline–online decomposition. The offline component requires just one adjoint wave propagation per sensor; using MFEM, we scale this part of the computation to the full El Capitan system (43,520 GPUs) with 92% weak parallel efficiency. Moreover, given real-time data, the online component exactly solves the Bayesian inverse and forecasting problems in 0.2 seconds on a modest GPU system, a ten-billion-fold speedup.

97 MATHEMATICS AND COMPUTING↗

Scalable Reduced Order Model with Discontinuous Galerkin Domain Decomposition

scaleupROM is a scalable, physics-constrained reduced order model (ROM). It aims to provide robust, accelerated physics predictions at extrapolated scales, based on the small, component-level data. This is implemented by combining projection-based ROM with discontinuous Galerkin domain decomposition, in the framework of MFEM and libROM. It currently supports the Poisson equation and Stokes flow equation, and more work is in progress toward general, nonlinear physics systems.

Chung, Seung Whan↗

Homotopy Solver

This software implements parallel versions of an interior-point solver, based on the publicly available ipopt solver. Here we have full control over the linear solver and our algorithm is fully parallel thus enabling scalability to large-scale optimization problems. This package also has a parallel implementation of a homotopy solver developed under the scalable methods for contact LDRD project 23-ERD-017. This solver is an mfem-based implementation of algorithm described in ``A filter trust-region Newton continuation method for nonlinear complementarity problems''. Cosmin G. Petra, Nai-Yuan Chiang, Jingyi Wang, Tucker Hartland, and Michael Puso (submitted), LLNL-JRNL-869761.

Hartland, Tucker [Lawrence Livermore National Labo↗

TensorFEM

The purpose of this research library is to demonstrate how to combine the BoBa tensor library and MFEM finite element library to enable tensorized finite element methods.

Guthrey, PiersonT [Lawrence Livermore National Lab↗

Ginkgo - A math library designed to accelerate Exascale Computing Project science applications

Large-scale simulations require efficient computation across the entire computing hierarchy. A challenge of the Exascale Computing Project (ECP) was to reconcile highly heterogeneous hardware with the myriad of applications that were required to run on these supercomputers. Mathematical software forms the backbone of almost all scientific applications, providing efficient abstractions and operations that are crucial to harness the performance of computing systems. Ginkgo is one such mathematical software library, nurtured by ECP, providing high-performance, user-friendly, and performance portable interfaces for applications in ECP and beyond. In this paper, we elaborate on Ginkgo’s philosophy of high-performance software that is sustainable, reproducible, and easy to use. We showcase the wide feature set of solvers and preconditioners available in Ginkgo and the central concepts involved in their design. We elaborate on four different ECP software integrations: MFEM, PeleLM + SUNDIALS, XGC, and ExaSGD that use Ginkgo to accelerate their science runs. Performance studies of different problems from these applications highlight the effectiveness of Ginkgo and the benefits incurred by these ECP applications.

Cojean, Terry↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗