Search NASA⌕ Search

SEARCH · Search NASA

Results for “performance gain”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

TunIO: An AI-powered Framework for Optimizing HPC I/O

I/O operations are a known performance bottleneck of HPC applications. To achieve good performance, users often employ an iterative multistage tuning process to find an optimal I/O stack configuration. However, an I/O stack contains multiple layers, such as high-level I/O libraries, I/O middleware, and parallel file systems, and each layer has many parameters. These parameters and layers are entangled and influenced by each other. The tuning process is time-consuming and complex. In this work, we present TunIO, an AI-powered I/O tuning framework that implements several techniques to balance the tuning cost and performance gain, including tuning the high-impact parameters first. Furthermore, TunIO analyzes the application source code to extract its I/O kernel while retaining all statements necessary to perform I/O. It utilizes a smart selection of high-impact configuration parameters of the given tuning objective. Finally, it uses a novel Reinforcement Learning (RL)-driven early stopping mechanism to balance the cost and performance gain. Experimental results show that TunIO leads to a reduction of up to ≈73% in tuning time while achieving the same performance gain when compared to H5Tuner. It achieves a significant performance gain/cost of 208.4 MBps/min (I/O bandwidth for each minute spent in tuning) over existing approaches under our testing.

Rajesh, Neeraj↗

Optimization of the light detection system of the ICARUS detector

The Short Baseline Neutrino (SBN) Program at Fermilab is designed to investigate short-baseline neutrino oscillations and test the hypothesis of sterile neutrinos, motivated by several experimental anomalies observed over the past decades. Within this program, the ICARUS experiment plays a key role. It employs the world’s largest Liquid Argon Time Projection Chamber (LArTPC) and serves as the farthest and most sensitive SBN detector for studying muon and electron neutrino oscillations. A crucial subsystem of the ICARUS detector is the Light Detection System (LDS), which captures the prompt scintillation light produced by neutrino interactions in the 600-ton active liquid Argon volume. This system provides precise timing information that is essential for event reconstruction, the trigger system, and cosmic background rejection. The LDS is composed of 360 Hamamatsu R5912-MOD 8-inch photomultiplier tubes (PMTs), operating under cryogenic conditions ($\sim 87 \ K$) inside the detector’s cryostats. During the detector’s operation at FNAL, a degradation in PMT gain has been observed, attributed to aging under low-temperature conditions. In collaboration with ICARUS teams from INFN Pavia and Catania, I developed an experimental setup to study the temperature-dependent behavior of the PMTs, performing gain measurements both at room temperature and down to $-70°C$ using a climatic chamber at INFN Catania. The results indicate that while the PMTs maintain stable gain at room temperature, a significant and permanent gain reduction occurs at low temperatures. Although $-70°C$ is still warmer than liquid Argon temperatures, the findings clearly demonstrate a gain-dependent performance degradation. The thesis also discusses mitigation strategies implemented in the ICARUS detector to address this issue and presents a simplified model to describe and simulate the observed behavior.

Saia, Clara [Catania U.] (ORCID:0009000464102417)↗

Quantum Computing and Visualization: A Disruptive Technological Change Ahead

The focus of this Visualization Viewpoints article is to provide some background on quantum computing (QC), to explore ideas related to how visualization helps in understanding QC, and examine how QC might be useful for visualization with the growth and maturation of both technologies in the future. In a quickly evolving technology landscape, QC is emerging as a promising pathway to overcome the growth limits in classical computing. In some cases, QC platforms offer the potential to vastly outperform the familiar classical computer by solving problems more quickly or that may be intractable on any known classical platform. As further performance gains for classical computing platforms are limited by diminishing Moore’s Law scaling, QC platforms might be viewed as a potential successor to the current field of exascale-class platforms. Importantly, while present-day QC hardware platforms are still limited in scale, the field of quantum computing is robust and rapidly advancing in terms of hardware capabilities, software environments for developing quantum algorithms, and educational programs for training the next generation of scientists and engineers. After a brief introduction to QC concepts, the focus of this article is to explore the interplay between the fields of visualization and QC. First, visualization has played a role in QC by providing the means to show representations of the quantum state of single-qubits in superposition states and multiple-qubits in entangled states. Second, there are a number of ways in which the field of visual data exploration and analysis may potentially benefit from this disruptive new technology though there are challenges going forward.

97 MATHEMATICS AND COMPUTING↗

Optimization of the light detection system of the ICARUS detector

The ICARUS detector, a key component of the Short Baseline Neutrino (SBN) Program at Fermi National Acelerator Laboratory (FNAL), is a 600-ton Liquid Argon Time Projection Chamber (LArTPC) equipped with a Light Detection System (LDS) that uses 360 Hamamatsu R5912-MOD 8-inch photomultiplier tubes (PMTs), specifically designed to operate under cryogenic conditions ($\sim 87 \ K$). These PMTs feed the trigger signal to the readout, improve the spatial and timing resolution of the events, and contribute to cosmic rays mitigation. During operation at FNAL, a progressive degradation in the PMT gain was observed. We developed an experimental setup to investigate the temperature dependence of PMT performance. Gain measurements were carried out from room temperature to $-70 ^\circ C$ using an environmental chamber. The results show that, while the PMTs exhibit stable performance at room temperature, a significant and irreversible reduction in gain emerges at lower temperatures. Al though $-70 ^\circ C$ remains above the liquid argon temperatures, the trend clearly reveals a gain-sensitive degradation mechanism. A simplified physical model was developed to reproduce and interpret the observed behavior. Based on these findings, a series of mitigation strategies were implemented in the ICARUS detector to preserve PMT performance and ensure reliable operation under cryogenic conditions.

Saia, C. [Catania Astrophys. Observ.]↗

Conceptual Design of the Moderator Test Station at the Spallation Neutron Source

We describe the Moderator Test Station (MTS), under construction at the Spallation Neutron Source. We will leverage the Beam Test Facility (BTF) at the Spallation Neutron Source to provide a moderator neutronics test stand with which we will verify the anticipated performance gains expected and required from the innovative moderator concepts central to the SNS Second Target Station (STS) and the future of the First Target Station (FTS). These concepts include large-volume para-hydrogen moderators and high-brightness tube moderators. The MTS will add to the existing BTF a proton beam chopper similar to that already used in the SNS at the RFQ exit (the so-called MEBT chopper), various proton beam transport components, a neutron-producing target, a cryogenic moderator test stand, a reflector-shielding assembly, and a performance assessment neutron beamline. The MTS will provide the ability to test a wide variety of moderators in a prototypic configuration, simultaneously measuring the neutronic performance of the moderator concept central to the anticipated STS gains and developing the instrumentation necessary to provide that performance in a production environment. The presentation will describe the MTS, its progress, and the way we optimize the prototypic nature of the moderators we will test.

Iverson, Erik B↗

LibraryX: A Framework for Cross-Library-Call Optimization

Scientific applications utilize performance libraries as a software engineering concept: these libraries encapsulate important and well-understood (mathematical) operations, allow for reuse, and are implemented and tuned by experts. Domain scientists then implement complex algorithms based on these domainspecific libraries. While individual library calls are optimized, larger performance gains across sequences of calls—sometimes spanning multiple libraries—are often unrealized, forcing a trade-off between performance and implementation complexity.To overcome this issue, we propose LibraryX, an approach and a system that allows for cross-library-call optimization even when library calls stem from multiple performance libraries. LibraryX annotates library calls with semantic information and optimizes entire directed acyclic graphs (DAGs) of calls dynamically using the SPIRAL code generation system. We demonstrate its effectiveness across a range of memory bound workloads, achieving significant speedups on Nvidia, AMD, and Intel accelerators compared to code using native libraries without cross-call optimization.

Rao, Sanil [Carnegie Mellon University,Department ↗

On the emerging potential of quantum annealing hardware for combinatorial optimization

Abstract Over the past decade, the usefulness of quantum annealing hardware for combinatorial optimization has been the subject of much debate. Thus far, experimental benchmarking studies have indicated that quantum annealing hardware does not provide an irrefutable performance gain over state-of-the-art optimization methods. However, as this hardware continues to evolve, each new iteration brings improved performance and warrants further benchmarking. To that end, this work conducts an optimization performance assessment of D-Wave Systems’ Advantage Performance Update computer, which can natively solve sparse unconstrained quadratic optimization problems with over 5,000 binary decision variables and 40,000 quadratic terms. We demonstrate that classes of contrived problems exist where this quantum annealer can provide run time benefits over a collection of established classical solution methods that represent the current state-of-the-art for benchmarking quantum annealing hardware. Although this work does not present strong evidence of an irrefutable performance benefit for this emerging optimization technology, it does exhibit encouraging progress, signaling the potential impacts on practical optimization tasks in the future.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Bifacial PV Module Energy Modeling Validation Study

The cost delta between monofacial and bifacial photovoltaic (PV) modules was decreasing in 2018 when this study was initially proposed, making bifacial modules an attractive offering. However, there were limited field studies available that could be considered utility-scale representative, and PV module manufacturers were stating large ranges for the expected energy yield improvements of bifacial technology. This generated uncertainties regarding bifacial performance gains and the actual LCOE (levelized cost of energy). There was also relatively low confidence in the industry’s energy modeling tools’ ability to accurately predict bifacial energy yield, meaning that early adopters of bifacial modules were not able to fully account for the increased energy yield in their energy and financial models. The intent of this project was to provide the solar industry with greater certainty on the energy yield gains of bifacial modules and provide higher confidence in the ability for different energy modeling software to model bifacial PV site performance. Achieving this would allow for bifacial technology to become bankable (i.e. accepted by Independent Engineers, site financiers and other PV site stakeholders), offering a step change in PV site performance that had not been seen since the widespread adoption of the single-axis tracker. Over the course of this study other economic factors including a significant tariff exemption for bifacial modules resulted in accelerated adoption of bifacial technology throughout the US utility scale solar segment. The number of early adopters increased rapidly, and with many investors accepting this module choice the question of bifacial bankability was seemingly answered faster than expected. However, PVEL’s study has still achieved significant accomplishments in demonstrating the accuracy of bifacial modeling across three different software platforms. The results of this study have shown that the mean biased error (MBE) between the field data and the predicted values from all three software platforms were aligned with a maximum MBE of 1.3% and a minimum MBE of -1.8%.

14 SOLAR ENERGY↗

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS↗

Performance evaluation of automated data-driven feature extraction and selection methods for practical and scalable building energy consumption prediction models

Here, this study quantifies the impact of automated feature engineering methods (feature extraction and selection) on the quality and accuracy of machine learning models that predict building energy consumption. The case study compares model performance for three main scenarios: baseline (no feature extraction and selection), feature extraction only, and feature extraction combined with feature selection (filter and/or wrapper methods) for fully trained machine learning models for 200 metered/sub-metered energy measurements across 118 real buildings. For consistency, the same machine learning model architecture (a black box deep learning neural network with probabilistic forecast output) was used for all scenarios. Based on results, all feature engineering methods provided noticeable prediction accuracy improvements (e.g., 29%-68% median prediction improvement) compared to baseline scenarios. However, in this application, feature selection methods provide little practical value due to their limited performance gains and high computational cost. Smarter algorithm development supported by better computational environments will be needed before feature selection methods can reliably and efficiently improve predictive model performance.

97 MATHEMATICS AND COMPUTING↗

Single and Double-Sided Jet Impingement Cooling for SiC-Based Power Modules

Efficient thermal management of power electronics systems is crucial for higher reliability. With the miniaturization of systems, high-loss-density electronics require cooling systems that can extract a large amount of heat. This study explored a liquid-jet-impingement-based direct substrate cooling system for single-sided and double-sided cooling to improve heat extraction efficiency and improve the power density by reducing the volume and mass. The cooling system was implemented for a SiC-based direct bonded copper substrate. Numerical simulations were performed to determine the effects of nozzle diameter, the number of nozzles, and nozzle array orientation on single-sided cooling and thermal performance gain over double-sided cooling. A novel manifold design was proposed that reduced the volume and mass of the manifold and still achieved the target power density. The performance of the proposed design was compared with the pin-fin-based cooling system used in the BMW I3 module, and a comparative analysis was done.

Barua, Himel↗

Achieving 20.1% Efficiency in Organic Solar Cells Through Interconnected Fibrillar Networks via Local Molecular Stacking

ABSTRACT The performance of organic solar cells (OSCs) is governed by how molecular packing evolves into interconnected networks that facilitate exciton dissociation and charge transport. Using an all‐small‐molecule blend DR3TSBDT:Y6 as a model system, we study how local molecular stacking evolves into performance‐relevant morphology during solvent vapor annealing (SVA) and subsequent thermal annealing (TA). SVA promotes end‐to‐end stacking of amorphous acceptors to form interconnected fibrils, while TA compacts inter‐fibril spacing without disrupting favorable local order. Such molecular‐to‐morphological refinements broaden light absorption, enhance charge transport, and markedly improve device efficiency. Extending this approach to additional blend systems (D18:Y6, D18:L8‐BO, and DR3TSBDT:L8‐BO) yields similar structural evolution and performance gains, with the D18:L8‐BO system achieving up to 20.10% PCE. Our study establishes control over local stacking in amorphous acceptors into fibrillar networks as a general and effective route to realize high‐performance OSCs.

Wu, Junying↗

Blade and Rim Seal Design of a First Stage High Pressure Turbine for a 300 MWe Supercritical CO2 Power Cycle

A first stage high-pressure turbine (HPT) blade is optimized for a 300 MWe supercritical CO2 (sCO2) power cycle using the surrogate-assisted genetic algorithm optimizer in Numeca FINE/Design 3D with objectives of increasing efficiency and decreasing heat load to the blade. The National Institute of Standards and Technology Reference Fluid Thermodynamic and Transport Properties Database (NIST REFPROP) [1] data for supercritical CO2 is formatted into tables of bicubic polynomial coefficients for use in condensable gas simulations in FINE/Turbo. Nearly 3000 unique shapes are evaluated via three-dimensional Reynolds Averaged Navier Stokes simulations, yielding increases in efficiency of up to 0.85 percentage points and decreases in heat load of 14%. A final blade, deemed the advanced blade, is chosen for future experimental analysis. Following this, a squealer tip optimization is performed on both the baseline and advanced blade designs. This optimization resulted in a performance gain of 1.25 points in efficiency and 15% reduction in tip heat load compared to the baseline flat tip design at the same clearance. In tandem, an optimization of the rotor-stator platform rim seal is performed using a parametrized geometry allowing for straight, meandering, and knife seal cavities. This multi-objective optimization focuses on decreasing the cooling mass flow and increasing the heat flux from the rotor and stator disk. The optimization resulted in cooling mass flow decreases of up to 26% while maintaining the average heat flux on the rim seal.

Tuite, Logan↗

Pressure Gain, Stability, and Operability of Methane/Syngas Based RDEs Under Steady and Transient Conditions (Final Project Report)

The scope of this work addresses key issues associated with losses associated with the detonation wave and other processes internal to the RDE operation, as well as it develops modeling tools for the evaluation of these losses and exhaust emissions in RDEs. The main challenge in studying RDEs is that RDE performance is highly reliant on the specifics of the design so much so that simple/canonical systems alone cannot provide useful engineering information, but practical RDE designs are sufficiently complex and involve extreme operational environments that detailed access either experimentally (laser diagnostics, for instance) or computationally (direct numerical simulations) are as yet to become practical. To overcome this challenge, we have conducted a combined experimental/simulation/analytical study investigating key phenomena that control the characteristics of operation of RDEs. As a result, the study has developed tools and methods that can be used to evaluate performance and design approaches using reduced-physics models, with the assumptions validated using detailed simulations, and the model prediction tested using experimental observations. The specific objectives of the research were: (1) Develop and demonstrate a low-loss fully axial injection concept, taking advantage of stratification effects to alter the detonation structure and position the wave favorably within the combustor; (2) Obtain stability and operability characteristics of an RDE across operating conditions to aid in the development of operability and performance rules for the operations of other systems; and (3) Develop quantitative metrics for performance gain as well as quantitative description of the loss mechanisms through a combination of diagnostics development, reduced-order modeling, and detailed simulations. The work conducted here has made contribution on design of low-loss inlets that has broad application within the power generation industry for use with pressure gain combustion. The operability and stability of different designs, while focusing on axial air inlet designs, has been analyzed. The effect of nozzle and injection conditions was studied. Models and simulations of exhaust emissions, focusing on NOx emission has been developed and used to investigate how operation of the RDE affect NOx production using Lagrangian analysis of RDE simulations. This work has built on previous programs, with the goal of further understanding operation of RDEs and elevate the readiness of design consideration. In addition, a suite of diagnostic and modeling tools have been developed to obtain quantitative metrics on performance based on measurements, which can be readily transferred to other experimental configurations.

08 HYDROGEN↗

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

Towards High-Performance AI4NP Applications on Modern GPU Platforms

The evolution of modern heterogeneous accelerators, such as GPUs, has significantly advanced the landscape of artificial intelligence (AI). There is a notable surge to adopt AI within the nuclear physics domain (AI4NP). While most AI4NP studies focus on feasibility analysis, our attention is directed towards evaluating their performance on contemporary GPUs that integrate tensor cores. We first benchmark the throughput of hyperparameterized multi-layer perceptron (MLP) models. We then examine the performance of an AI4NP application: Hydra. We assess the performance gain and accuracy loss caused by the tensor cores for low-precision floating-point operations. Our experiments encompass the PyTorch and TensorFlow Keras frameworks on NVIDIA’s T4 and A100 GPUs. We explore the behavior of different GPU hardware platforms and AI software tools. This study can be a valuable resource for guiding the performance optimization of larger-scale deployments of AI4NP applications.

Mei, Xinxin↗

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗

A Low-Rank QTT-based Finite Element Method for Elasticity Problems

We present an efficient and robust numerical algorithm for solving the linear elasticity problem that combines the Quantized Tensor Train format and a domain partitioning strategy. This approach makes it possible to solve the linear elasticity problem on a computational domain that is more general than a square. By integrating Z-ordering and subdomain concatenation, our method substantially decreases memory usage and achieves a notable reduction in rank compared to established Finite Element implementations like the FEniCS platform. This efficiency is maintained while still guaranteeing exponential convergence with respect to the number of degrees of freedom. This performance gain, however, requires a fundamental rethinking of how core finite element operations are implemented. This includes changes to mesh discretization, node and degree of freedom ordering, stiffness matrix and internal nodal force assembly, and the execution of algebraic matrix-vector operations. In this work, we discuss all these aspects in detail and assess the method’s performance in the numerical approximation of three representative test cases.

97 MATHEMATICS AND COMPUTING↗