Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Large-scale real-time signal processing in physics experiments: the ALICE TPC FPGA pipeline

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB s -1 of raw detector data. This requirement is met by a custom FPGA-based processing pipeline that performs the complete front-end data treatment fully in-stream, including common-mode correction, pedestal subtraction, ion-tail filtering, zero suppression, and dense data packing. A central element of the design is a highly parallel common-mode correction algorithm operating directly on the streaming data. It robustly identifies signal-free readout channels on a time-bin basis and applies pad-dependent scaling to compensate for local variations in capacitive coupling in the GEM readout. In combination with pedestal subtraction and ion-tail filtering, this enables accurate baseline restoration under extreme high-occupancy conditions, preventing signal loss while efficiently suppressing noise prior to zero suppression. The pipeline operates continuously at the full detector bandwidth and reduces the raw input rate of approximately 3 TB s -1 to about 900 GBps for Pb-Pb collisions at the target interaction rate. Overall, it represents a large-scale FPGA-based real-time signal-processing implementation for high-energy physics detector readout.

Digital signal processing (DSP)↗

Optimizing Desalination Operations for Energy Flexibility

Despite the value of energy optimization in desalination processes, modeling dynamic operations for monthly billing periods has remained a computational challenge. This work proposes a framework for energy flexibility optimization, which includes new modeling features for independent operation of parallel skids, start-up delays associated with chemical stabilization, the consideration of industrial energy tariff structures, and inclusion of hourly electrical carbon intensities. This is done using a modular and computationally efficient formulation that guarantees a globally optimal solution with standard optimization solvers. In this study, the approach is demonstrated in two distinct case studies: a seawater desalination plant in Santa Barbara, CA, and an indirect potable reuse facility in San Jose, CA. Trends predicted from the model are validated against operational facility measurements from a demand response shutdown event. Preliminary results show that optimizing energy flexibility can result in 18.51% monthly cost savings over energy efficiency-optimized operation. The value extracted from a facility-wide shutdown during peak electricity price hours is hampered by start-up delays in post-treatment chemical stabilization. In cases in which a facility does not have much excess capacity, using a flow equalization tank or operating over a wide recovery range may be cost-effective.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermodynamic Limits of Redox-Based Thermochemical Processes (REDOTHERM)

Solar thermochemical fuel production is a potential pathway for the production of sustain liquid drop-in fuels, which can help decarbonize the aviation and maritime sectors. In an attempt to analyze the commercial viability of this technology, several studies have been conducted, including system and technoeconomic analysis (TEA) modeling. However, most studies to date simply assume a given redox reactor efficiency, which is significantly higher than demonstrated values to date. While it is widely recognized that utilizing a counter-current flow (CF) configuration could increase the redox reactor efficiency, an over-simplification in the thermodynamic modeling may lead to unphysical results which has been included in multiple publications. The fact that the solar redox reactor is the least developed component in the process chain makes it hard to identify technology gaps and evaluate pathways to deployment at scale using this approach. In this work, a thermodynamic model for a moving oxide system has been developed, in a general form that allows to analyze the system for different redox-active materials, under a wide range of operating conditions, for both parallel and countercurrent flows. The model capabilites are demonstrated, and the model's code will be shared as an open-source on GitHub in the next few months.

chemical looping↗

Low power on-chip data transmission for wafer-scale monolithic active pixel sensors

Here, this paper details the implementation of the digital pulse shaping subsystem within the Backbone Transmission Line Encoding (BTLE) driver, a low-power, long-distance on-chip data transmission solution designed in a 65 nm CMOS process. Digital pulse shaping is critical for minimizing inter-symbol interference (ISI) caused by bandwidth limitations of on-chip interconnects, especially in wafer-scale monolithic active pixel sensors (MAPS). A duobinary encoder coupled with a parallelized polyphase finite impulse response (FIR) filter is used for efficient shaping of the transmitted signal spectrum. This reconfigurable architecture achieves reliable 160 Mb/s data transfer over a 10 cm on-chip link, as validated by simulations demonstrating low power consumption (FoM 37.3 fJ/bit/mm of transmission line length) and effective ISI mitigation.

47 OTHER INSTRUMENTATION↗

Development of Steady-State and Dynamic Mass and Energy Constrained Neural Networks for Distributed Chemical Systems Using Noisy Transient Data

The paper presents the development of algorithms for mass and energy constrained neural network models that can exactly conserve the overall mass and energy of distributed chemical process systems, even though the noisy transient data used for optimal model training violate the same. In contrast to approximately satisfying mass and energy balance constraints of a system by soft penalization of objective function, algorithms have been developed for solving equality-constrained nonlinear optimization problems, thus providing the guarantee of exactly satisfying the system mass and energy conservation laws. For developing dynamic mass-energy constrained network models for distributed systems, hybrid series and parallel dynamic-static neural networks have been leveraged. The developed algorithms for solving both the training and forward problems are validated using both steady-state and dynamic data in the presence of various noise characteristics. The developed data-driven algorithms are flexible to exactly satisfy mass and energy balance constraints for dynamic chemical processes if the system holdup information is available. The proposed network structures and algorithms are applied to the development of data-driven lumped and distributed models of an adiabatic superheater/reheater system, a nonisothermal continuous stirred tank reactor, as well as an electrically heated plug-flow reactor system where one form of energy gets transformed to another. It has been observed that the mass-energy constrained neural networks yield a root mean squared error of <1% with respect to the system truth for the case studies evaluated in this work.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Model-based, in-situ, non-destructive qualification and certification of parts made by autonomous additive manufacturing

To address the significant productivity challenges associated with the qualification and certification (Q&C) tasks of additively manufactured (AM) parts, which have traditionally relied on rigorous post‐build inspection and testing, we propose an integrated framework that combines model‐based qualification and certification (MBQ&C) with autonomous additive manufacturing (AAM). MBQ&C employs high‐fidelity predictive models, developed within the Integrated Computational Materials Engineering (ICME) paradigm, to simulate process–structure–property–performance relationships for assessing a part’s fitness for use. Since predictive models are commonly machine learning (ML)-based or reduced-order surrogates of validated physics models, they run efficiently, enabling timely inference. In parallel, the self-driving AAM utilises ML-based adaptive, closed‐loop control strategies to avoid, mitigate, or repair defects and anomalies during fabrication, thereby increasing the likelihood of producing acceptable parts. A key feature of the combined AAM-MBQ&C framework is that predictive models explicitly incorporate defects or anomalies that persist after the build, using instance-specific data captured via in-situ sensing. This customisation enables a build‐specific assessment of fitness for use, rather than relying on nominal or generic parameters. Such individualised evaluation provides a robust basis for Q&C-related acceptance decisions relating to each build. Additionally, the rapid solution capabilities of ML or reduced-order models enable the determination of a part’s suitability for service shortly after build completion. As the framework matures, it has the potential to substantially reduce reliance on conventional point‐design approaches—such as time‐consuming post‐build computed tomography scanning and costly destructive testing. Thus, the AAM-MBQ&C framework represents a transformative, scalable strategy for quality assurance of AM components, as parts produced within a stable, validated, and certified envelope can be certified with reduced testing. Key benefits include: (1) significant gains in Q&C productivity through efficient, model-centric assessment; (2) performance-based classification of defects into critical and non-critical categories; (3) the ability to predict potential deviations in the performance of parts affected by real-time, adaptive process control interventions relative to those produced under a certified process, and (4) the enabling of virtual Q&C for service environments that are difficult, hazardous, or impractical to access or reproduce experimentally. Collectively, these capabilities strengthen the business case for AM, particularly for high‐consequence and mission‐critical applications. Finally, although this work focuses on powder-based AM, the proposed techniques could be extended to AM processes employing alternative feedstock forms.

Gunasegaram, Dayalan↗

Effects of wave damping and finite perpendicular scale on three-dimensional Alfvén wave parametric decay in low-beta plasmas

Shear Alfvén wave parametric decay instability (PDI) provides a potential path toward significant wave dissipation and plasma heating. However, fundamental questions regarding how PDI is excited in a realistic three-dimensional (3D) open system and how the finite perpendicular wave scale—as found in both laboratory and space plasmas—affects the excitation remain poorly understood. Here, we present the first 3D, open-boundary, hybrid kinetic-fluid simulations of kinetic Alfvén wave PDI in low-beta plasmas. Key findings are that the PDI excitation is strongly limited by the wave damping present, including electron–ion collisional damping (represented by a constant resistivity) and geometrical attenuation associated with the finite-scale Alfvén wave, and ion Landau damping of the child acoustic wave. The perpendicular wave scale alone, however, plays no discernible role: waves of different perpendicular scales exhibit similar instability excitation as long as the magnitude of the parallel ponderomotive force remains unchanged. These findings are corroborated by theoretical analysis and estimates. This new understanding of 3D kinetic Alfvén wave PDI physics is essential for laboratory study of the basic plasma process and may also aid future evaluation of the relevance/role of PDI in low-beta space plasma.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Cryogenic readout integrated circuit with analog pile-up and in-Pixel ADC for high frame rate Skipper CCD-in-CMOS Sensors

The Skipper CCD-in-CMOS Parallel Read-Out Circuit V2 (SPROCKET2) is designed to enable high frame rate readout of Skipper CCD-in-CMOS image sensors. The SPROCKET2 pixel is fabricated in a 65 nm CMOS process and occupies a 60$\mu$m $\times$ 60$\mu$m footprint. SPROCKET2 is intended to be heterogeneously integrated with a pixelated Skipper CCD-in-CMOS sensor, such that one readout pixel is connected to a multiplexed array of 16 active image sensor pixels, to match their spatial geometry. Our design benefits from the Skipper CCD-in-CMOS sensor's non-destructive readout capability to achieve exceptionally low noise through multi-sampling and averaging while optimizing for total power consumption. The pixel readout utilizes correlated double sampling to minimize 1/f noise and includes "pile-up" of ten successive samples in the analog domain before digitizing at a rate of 66.7 ksps. Measurement results of in-pixel serial SAR ADC show DNL and INL of ~0. 44 LSB and 0.58 LBS respectively. A large area array of 20,000 SPROCKET2 ADC pixels (multiplexed 1:16 to 320,000 sensor pixels) is currently under test. By reading out data over a 10 Gbps optical link, this pixel design enables a frame rate of $\sim$ 4 kfps for large sensing areas with minimal sensing deadtime. In the highest gain mode, the pixelated ADC has an input-referred resolution of 10$\mu$V with a simulated power consumption of 50$\mu$W. The pixel operates with constant current draw to minimize power-rail crosstalk.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Self-Trapped-Exciton Radiative Recombination in β–Ga 2 O 3 : Impact of Two Concurrent Nonradiative Auger Processes

The peculiarities of radiative and nonradiative processes associated with self-trapped intrinsic eXcitons in the excited β-Ga 2 O 3 crystals are studied via time-resolved techniques of induced absorption, transient grating, and photoluminescence (PL) at room temperature. The excitation above the bandgap is produced by laser pulses with linear light polarization parallel and orthogonal in the (–201) and (001) planes. We elucidate that the nonradiative recombination rate occurring in the eXciton prevails over its radiative emission rate in a wide range of free carrier concentration composed of excited and equilibrium electrons. Hence, the nonradiative recombination has no effect on the strong anisotropy and the shape of the eXciton emission band. However, we find out that the conventional ABC model of electron effective lifetime is insufficient for explanation of the excitation dependences. Inclusion of two nonradiative Auger mechanisms in a modified ABC formula provides excellent agreement of these dependences. We conclude that the trap-assisted Auger process is in proportion to the free electron density with coefficient B = 1.1 × 10 –11 cm 3 /s and appears at low/intermediate excitation, while the triple-particle Auger process is in proportion to Δn 2 with coefficient C = 8 × 10 –30 cm 6 /s and appears at high excitation conditions. The transition between two Auger mechanisms is accompanied by a rise of the eXciton diffusivity in preferred crystallographic directions where the radiative PL intensity is maximal. The diffusion length LD in these directions can reach values ~300 nm, but, at high excitations, LD becomes limited by Auger lifetimes. These findings pave the way for the implementation of self-trapped eXcitons into specific optoelectronic devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Controlled patterning of crystalline domains by frontal polymerization

Materials with hierarchical architectures that combine soft and hard material domains with coalesced interfaces possess superior properties compared with their homogeneous counterparts. These architectures in synthetic materials have been achieved through deterministic manufacturing strategies such as 3D printing, which require an a priori design and active intervention throughout the process to achieve architectures spanning multiple length scales. Here we harness frontal polymerization spin mode dynamics to autonomously fabricate patterned crystalline domains in poly(cyclooctadiene) with multiscale organization. This rapid, dissipative processing method leads to the formation of amorphous and semi-crystalline domains emerging from the internal interfaces generated between the solid polymer and the propagating cure front. The size, spacing and arrangement of the domains are controlled by the interplay between the reaction kinetics, thermochemistry and boundary conditions. Small perturbations in the fabrication conditions reproducibly lead to remarkable changes in the patterned microstructure and the resulting strength, elastic modulus and toughness of the polymer. Furthermore, this ability to control mechanical properties and performance solely through the initial conditions and the mode of front propagation represents a marked advancement in the design and manufacturing of advanced multiscale materials. Drawing inspiration from biological systems in which structural complexity develops through dissipative reaction–diffusion processes, this study explores a transformative synthetic manufacturing strategy aimed at harnessing the principles underpinning morphogenic growth, unlocking new avenues for advanced materials design and fabrication. Synthetic coupled reaction-transport processes offer a versatile yet relatively underexplored method to manipulate the spatial attributes of synthetic materials10. Here we introduce an innovative manufacturing approach based on frontal ring-opening metathesis polymerization (FROMP) that draws parallels with morphogenic growth and development, enabling the formation of patterned microstructures within polymeric materials.

36 MATERIALS SCIENCE↗

Towards the First High-Q Treatments for the FCC 800 MHz 5-Cell Elliptical Cavities

Development towards the various realizations of the FCC machine requires optimization of sub-GHz elliptical cavities for high-gradient and high-Q operation, both in pulsed and CW mode, for application in the booster and collider portions. Previous development work validated the proposed 800 MHz 5-cell elliptical RF design, showing reasonable performance after EP treatment. However, the stringent high-Q (3.8e+10) and high-gradient (24 MV/m) goals of the FCC machine cavities will require further development, relying on advanced surface processing techniques developed at 1.3 GHz, such as medium-temperature furnace baking. We describe the development and preparation of 1- and 5- cell 800 MHz cavities for the high-Q program. In parallel, we discuss the design progress and strategies for integrating the 800 MHz cavities into cryomodules to be implemented in both the booster and collider rings.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

BM3DORNL

BM3DORNL is a high-performance, open-source library for removing streak and ring artifacts from computed-tomography (CT) data, developed for neutron imaging at Oak Ridge National Laboratory's Spallation Neutron Source (VENUS beamline) and applicable to X-ray CT as well. Ring artifacts — concentric rings in reconstructed slices caused by detector pixel-to-pixel response non-uniformities — appear as vertical streaks in the sinogram and degrade both image quality and quantitative analysis. BM3DORNL operates in the sinogram domain using an adaptation of the BM3D (block-matching and 3D collaborative filtering) algorithm (Dabov et al., 2007). It provides a dedicated streak-removal mode, a true multi-scale BM3D variant (after Mäkinen et al., 2021) that suppresses wide streaks single-scale methods miss, and an alternative Fourier–SVD method (~2.6× faster) combining FFT-based energy detection with rank-1 SVD. The computationally intensive core is implemented in Rust with parallel (Rayon) block matching, integral-image pre-screening, and optimized transforms, and is exposed through a simple Python API (with an optional GUI) so it integrates directly into existing tomography reconstruction pipelines. It processes both 2D sinograms and 3D sinogram stacks, is pip-installable for Linux and macOS, and is documented at https://bm3dornl.readthedocs.io.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

PIAFS: A 2D nonlinear hydrodynamics code to model gaseous optics

The survivability of final optics is expected to be a major challenge for all future inertial fusion energy concepts. Due to their higher damage threshold, gaseous optics have been identified as a promising solution to this problem. Gaseous optics can be created through the photoabsorption of spatially modulated UV light, which induces various chemical processes that heat the gas. This heating leads to a pressure perturbation, which in turn launches a density perturbation that can imprint a refractive index modulation such as a grating. In this article, we introduce a parallel C/C++ code to simulate gaseous optics. PIAFS2D is a high-order conservative finite-difference code to solve the compressible Navier–Stokes equations along with the photochemical heating sources on Cartesian grids. The simulations are validated by the linear theory derived in a previous paper [Michel et al., Phys. Rev. Appl. 22, 024014 (2024)]. For larger perturbations, the behavior of the system—particularly the evolution of the generated acoustic wave—demonstrates strong nonlinearity. PIAFS2D allows the study of nonlinear behaviors and can be used for the design of high-efficiency gaseous optics elements in realistic experimental conditions.

Oudin, A. [Lawrence Livermore National Laboratory ↗

Sustainable aviation fuels from biomass and biowaste via bio- and chemo-catalytic conversion: Catalysis, process challenges, and opportunities

Sustainable aviation fuel (SAF) production from biomass and biowaste streams is an attractive option for decarbonizing the aviation sector, one of the most-difficult-to-electrify transportation sectors. Despite ongoing commercialization efforts using ASTM-certified pathways (e.g., lipid conversion, Fischer-Tropsch synthesis), production capacities are still inadequate due to limited feedstock supply and high production costs. New conversion technologies that utilize lignocellulosic feedstocks are needed to meet these challenges and satisfy the rapidly growing market. Combining bio- and chemo-catalytic approaches can leverage advantages from both methods, i.e., high product selectivity via biological conversion, and the capability to build C-C chains more efficiently via chemical catalysis. Herein, conversion routes, catalysis, and processes for such pathways are discussed, while key challenges and meaningful R&D opportunities are identified to guide future research activities in the space. Bio and chemo-catalytic conversion primarily utilize the carbohydrate fraction of lignocellulose, leaving lignin as a waste product. This makes lignin conversion to SAF critical in order to utilize whole biomass, thereby lowering overall production costs while maximizing carbon efficiencies. Thus, lignin valorization strategies are also reviewed herein with vital research areas identified, such as facile lignin depolymerization approaches, highly integrated conversion systems, novel process configurations, and catalysts for the selective cleavage of aryl C–O bonds. The potential efficiency improvements available via integrated conversion steps, such as combined biological and chemo-catalytic routes, along with the use of different parallel pathways, are identified as key to producing all components of a cost-effective, 100% SAF.

09 BIOMASS FUELS↗

The development and applications of multidimensional biomolecular spectroscopy illustrated by photosynthetic light harvesting

The parallel and synergistic developments of atomic resolution structural information, new spectroscopic methods, their underpinning formalism, and the application of sophisticated theoretical methods have led to a step function change in our understanding of photosynthetic light harvesting, the process by which photosynthetic organisms collect solar energy and supply it to their reaction centers to initiate the chemistry of photosynthesis. The new spectroscopic methods, in particular multidimensional spectroscopies, have enabled a transition from recording rates of processes to focusing on mechanism. We discuss two ultrafast spectroscopies – two-dimensional electronic spectroscopy and two-dimensional electronic-vibrational spectroscopy – and illustrate their development through the lens of photosynthetic light harvesting. Both spectroscopies provide enhanced spectral resolution and, in different ways, reveal pathways of energy flow and coherent oscillations which relate to the quantum mechanical mixing of, for example, electronic excitations (excitons) and nuclear motions. The new types of information present in these spectra provoked the application of sophisticated quantum dynamical theories to describe the temporal evolution of the spectra and provide new questions for experimental investigation. While multidimensional spectroscopies have applications in many other areas of science, we feel that the investigation of photosynthetic light harvesting has had the largest influence on the development of spectroscopic and theoretical methods for the study of quantum dynamics in biology, hence the focus of this review. We conclude with key questions for the next decade of this review.

59 BASIC BIOLOGICAL SCIENCES↗

A Performance and Energy Study of GPU-Resident Preconditioners for Conjugate Gradient Solvers: In the Context of Existing and Novel Approaches

Optimizing a particular subprogram out of the set of Basic (sparse) Linear Algebra Subprograms (BLAS) for a given architecture is a common topic of research. In applications, however, these BLAS functions rarely appear in isolation; usually, many of them are used together, in various combinations and with varying inputs. As the need to solve a large, sparse linear system is ubiquitous throughout HPC applications, linear solvers constitute a realistic, sufficiently complex and well-defined representative use case for composite BLAS routines. To this end, based on a representative set of matrices drawn from a diverse set of fields, we present a framework to study, from the performance and energy perspective, the efficacy of GPU- resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process. The development of this preconditioner was motivated by solving large graph Laplacian linear systems, for which the existing preconditioners either perform slow on GPU-based platforms or are not applicable. We compare the performance of these preconditioners on different hardware accelerator architectures, i.e., AMD MI250X, MI100, Nvidia A100, V100, and Jetson. Our experiments reveal performance trade-offs and provide information on how to select the best strategy for the given linear system, dictated by its properties, and the platform of interest. We demonstrate the application of our novel preconditioner for solving CG and graph Laplacian systems. Overall, the framework can be utilized as a benchmark to guide informed decisions in choosing a specific preconditioner, i.e., whether it is better to rely on the performance of a triangular solver or on the performance of sparse matrix-vector product. Finally, by considering power consumption to solve the linear systems, we report the energy footprint for the solvers.

Preconditioned Conjugate Gradient, GPUs, iterative↗