Search NASA⌕ Search

SEARCH · Search NASA

Results for “source separation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Simulation-Based Validation of An Open-Source, Scalable Framework for Building Energy Management in Small and Medium-Sized Commercial Buildings

Abstract: Small and medium-sized commercial buildings (SMCBs) represent 94% of U.S. commercial buildings but encounter substantial obstacles in adopting Building Energy Management (BEM) systems. Current approaches exhibit fundamental limitations: vendor-specific API platforms restrict interoperability through proprietary ecosystems; commercial automation software demands extensive technical expertise and licensing costs; open-source IoT solutions lack native support for building automation protocols and semantic models. This paper introduces a configuration-driven web interface framework addressing the gap between smart device advancements and accessible BEM software infrastructure for SMCBs. The framework leverages VOLTTRON middleware integrated with an automated converter that processes unified YAML configurations into heterogeneous system files, reducing required configuration artifacts from six separate files to a single unified specification. The system architecture enables vendor-agnostic operation through BACnet and Modbus protocols while supporting semantic building model integration via automated Brick Schema parsing. Configuration-driven interfaces automatically adapt to diverse HVAC types without custom development. Simulation-based validation using BOPTEST demonstrates automatic interface generation between fan coil and hydronic systems, with the automated converter successfully generating all platform-specific outputs from the single YAML input. The result demonstrates the framework's capability to streamline BEM system deployment through reduced configuration complexity. This work bridges simulation capabilities with operational deployment, demonstrating how virtual testbeds validate generalizable software frameworks for real-world building automation.

Chung, Jihoon [ORNL] (ORCID:0000000184880815)↗

High Temperature Test Facility At Scale Testing, Validation, and Demonstration

Theses slides will be used to brief the Tribal DOE and Council Chairman for the project listed below. The High Temperature Test Facility (HTTF) is a demonstration-scale high temperature electrolysis (HTE) system that provides the ability to produce, process, store, and dispense electrolytically produced hydrogen. The system will be designed to accommodate electrolysis systems from various industry partners. The HTTF will be supplied up to 10 MW of power for electrolysis at 5 separate test article locations, each up to 2MW. The system will include post-processing and compression, storage, and capabilities to provide hydrogen for consumption by future end users. The HTTF will also provide the ability to test and demonstrate an HTE system in a configuration that would simulate the use of a nuclear generated steam supply for the high temperature heat source. The HTTF will be located at the Central Facilities Area (CFA) adjacent to CFA-686 on the INL site, 45 miles west of Idaho Falls.

08 HYDROGEN↗

CMPO-Functionalized Silica Sorbents for pH-Tunable Separation and Enrichment of Rare-Earth Elements from Environmental Matrices

Rare-earth elements (REEs) are crucial in many applications, yet mutual separation is challenging due to their similar chemical behavior. Octylphenyl- N,N-diisobutyl carbamoyl methyl phosphine oxide (CMPO) is an organophosphorus ligand originally developed for extracting actinides and lanthanides from spent nuclear fuel. Here, we report a pH-tunable CMPOfunctionalized silica sorbent for selective REE separation from complex aqueous matrices. A CMPO-associated silica gel sorbent was synthesized and characterized by Brunauer−Emmett−Teller (BET) surface area, scanning electron microscopy, and X-ray photoelectron spectroscopy to confirm the surface functionalization and binding behavior. Sorbent performance was evaluated by using a synthetic 46- element solution and a real phosphate rock fertilizer leachate. Notably, REEs were successfully eluted with ultrapure water, demonstrating reversible desorption controlled by pH adjustment. Packed-bed column studies increased the REE mass fraction from 3.6% to 64% (20-fold enrichment), with up to 30-fold enrichment of neodymium. The adsorption process follows the Langmuir isotherm behavior and follows pseudo-second-order kinetics. The uptake capacity of 1 μmol of REEs per 4.2 μmol of CMPO supports the formation of a predominantly 4:1 ligand:rare earth element(III) pseudocomplex. These results demonstrate CMPO-functionalized silica as a selective, water-elutable, and low-chemical-input platform for sustainable REE recovery from environmental and industrial sources.

chelating ligands↗

Systematic Study of the Self-Renormalized Nucleon Gluon PDF in Large-Momentum Effective Theory

We present a systematic study of the nucleon gluon parton distribution function (PDF) using the self-renormalized large-momentum effective theory (LaMET) approach in lattice QCD. This work extends previous gluon-PDF extractions by performing a detailed analysis of key systematic effects, including gauge-link smearing, lattice spacing, pion mass, and nucleon boost momentum. The self-renormalization framework mitigates ultraviolet divergences associated with Wilson-line self-energy and renormalon contributions by combining lattice matrix elements with perturbative short-distance information, thereby preserving the correct infrared structure. Calculations are performed on $N_f=2+1+1$ HISQ ensembles generated by the MILC Collaboration at three lattice spacings and two pion masses, with boosted nucleon states reaching momenta up to 2.2~GeV. We determine renormalization factors from zero-momentum matrix elements and apply hybrid renormalization to suppress discretization artifacts. After extrapolating large-separation behavior and performing Fourier transforms, we reconstruct quasi-PDFs and match them to lightcone PDFs using next-to-leading order Wilson coefficients. Our results demonstrate that smearing and lattice-spacing effects are under control, and pion-mass and lattice-spacing dependence is mild relative to the current $O(10^6)$ statistics; however, momentum dependence remains a significant source of uncertainty. Future work including even larger boost momenta will be essential to reduce systematics in lattice determinations of the gluon PDF and to advance toward precision QCD phenomenology at the LHC and the future Electron-Ion Collider.

FOS: Physical sciences↗

A case study in contrastive learning information combination: Application to technical forensics of additive manufacturing filament source identification

Combination of information from disparate data sources into a single decision is a core challenge in many fields, including the field of technical forensics. Technical forensics (TF) utilizes technical characterization of questioned samples to determine properties of that sample; these properties are then used to infer information of forensic interest, such as provenance, age, or attribution. TF is utilized in traditional forensic applications, such as the attribution of material fragments from an explosive, and in nuclear forensic applications, such as the attribution of actinides which have been interdicted out of regulatory control. The challenge of combining information from disparate sources, described alternately by many terms including “Data Fusion” and “Data Integration”, is exacerbated in the technical forensics domain due to at least two factors: the challenge of interpreting each information source singularly, and the relatively small data set sizes available. Extensive literature exists attempting to combine technical forensics information sources, both in manual and automated processes. These attempts are often bespoke to the specific information sources (such as the bi-, tri-, or quad-isotope chart (Moody, Grant, and Hutcheon 2005)), with some emerging examples of simple early- and late- fusion (, respectively). Simultaneous to the information combination efforts described in the previous paragraph, the field of natural language processing attempted (and largely succeeded) in combining information from multiple non-technical information sources. The ecosystem of “multi-modal” language models, which can take text and images as input, and generate text and images as output, became large and diverse by 2025 (Khan et al. 2025). In a generalized sense, many of these methods are trained by learning neural networks which can convert raw text or images into a vector of numbers describing the text or image, hereafter called “embeddings” and the neural networks performing the conversion are called “embedders”. By using a separate embedder for text and images, finding coincident text and images (such as images with their captions), and optimizing the parameters of the embedders such that the embeddings for the text and the image are similar, the field has found a bridge between text and images (Girdhar et al. 2023). It is the contention of the authors of this report that this insight is not limited to text and images but instead can be extended to any modality which can be found coincidently. The subject of the rest of this report is the application of this method to example multi-modal technical forensic data. Some details about the data used in this report are not appropriate for this report, and are included in a companion report (PNNL-38669).

36 MATERIALS SCIENCE↗

Development and Implementation of a New AI-Based Tool to Support Fast Reactor Software Model Generation and Validation

This report summarizes FY26 work to develop Maggie, an artificial intelligence-based assistant designed to support software model generation and validation activities for fast reactor analysis codes. The project established a modular, code-agnostic software architecture that separates reusable agent capabilities from code-specific knowledge and tools, with initial implementation focused on the FRP-supported fast reactor safety analysis code SAS4A/SASSYS1 (SAS). A curated SAS-specific knowledge base was assembled from the code manual, training materials, historical analysis reports, and representative input files, and was integrated through retrieval-augmented generation to ground Maggie’s responses in authoritative sources. Maggie was deployed on the internal Argonne network, where it demonstrated practical user-facing capability as a chatbot for answering natural language questions about SAS and retrieving relevant technical information. Demonstration cases also showed that Maggie can generate useful snippets of SAS input for selected modeling tasks, while highlighting current limitations in reliability and consistency for more complex input generation tasks. Overall, the FY26 effort established the technical foundation for an AI-assisted capability intended to improve the efficiency, consistency, and accessibility of fast reactor software model development at Argonne and, with further improvements, to support eventual use by the broader fast reactor community, including industry users of FRP-supported analysis tools.

Thomas, Rachel [Argonne National Laboratory (ANL),↗

Sensitivity analysis of thermal contact conductance modeling to inform MiniFuel irradiation capsule designs

The MiniFuel irradiation platform has been developed by Oak Ridge National Laboratory as a flexible, high-throughput separate effects testing capability within the High Flux Isotope Reactor (HFIR). Finite element thermal models are relied upon to design MiniFuel experiments to achieve a specific time-averaged irradiation temperature for experimental objectives. A previous study identified that uncertainty in the component heat generation rates and thermal contact conductance (TCC) model are the most significant contributors to predicted fuel temperature variance. To address both sources of uncertainty, this work performs sensitivity analysis on the TCC model to identify high-impact, high-uncertainty parameters that contribute to fuel temperature variance. The TCC model is analyzed in increasing detail, first using a standalone Python code, then again after coupling Python to the BISON fuel performance code. Furthermore, the parameters with the largest contributions to fuel temperature variance which can be reduced through design changes are identified as the initial subcapsule gas pressure, contact pressure between the fuel and dish, and the effective surface roughness of the interface. A set of design recommendations for future capsule designs has been established and applied to reduce the previously quantified average fuel temperature uncertainty ranges of ± 40 °C in the HFIR vertical experiment facilities (VXF) and ± 80 °C in the removable beryllium (RB) reflector to approximately ± 32 °C and ± 53 °C, respectively. This equates to a 21 % and 33 % reduction in the uncertainty range of the average fuel temperature for VXF and RB, respectively.

BISON↗

Photosynthetic biohybrid systems for solar fuels catalysis

Photosynthetic reaction center (RC) proteins are finely tuned molecular systems optimized for solar energy conversion. RCs effectively capture and convert sunlight with near unity quantum efficiency utilizing light-induced directional electron transfer through a series of molecular cofactors embedded within the protein core to generate a long-lived charge separated state with a useable electrochemical potential. Of current interest are new strategies that couple RC chemistry to the direct synthesis of energy-rich compounds. This Feature Article highlights recent work from our lab on RC and RC-inspired hybrid systems that capture the Sun's energy and convert it to chemical energy in the form of H2, a carbon-neutral energy source derived from water. Further, biohybrids made from the Photosystem I (PSI) RC are among the best photocatalytic H2-producing protein hybrids to date. Targeted self-assembly strategies that couple abiotic catalysts to PSI translate to catalyst incorporation at intrinsic PSI sites within thylakoid membranes to achieve complete solar water-splitting systems. RC-inspired biohybrids interface synthetic photosensitizers and molecular catalysts with small proteins to create photocatalytic systems and enable the spectroscopic discernment of the structural features and electron transfer processes that underpin solar-driven proton reduction. In total, these studies showcase the incredible scientific opportunities photosynthetic biohybrid research provides for harnessing the optimal qualities of both artificial and natural photosynthetic systems and developing materials that capture, convert, and store solar energy as a fuel.

14 SOLAR ENERGY↗

Understanding the Mn dissolution mechanism in rock salt-type Li 4 Mn 2 O 5 cathodes

For the first time, a detailed exploration of Mn dissolution in disordered rock salt (DRX) Li 4 Mn 2 O 5 is presented. Herein, we apply a suite of synchrotron and lab scale X-ray techniques to both the cathode and the separator harvested from pristine, charged, or cycled lithium half-cells containing the disordered rock salt (DRX) material Li 4 Mn 2 O 5 , in order to understand Mn dissolution processes throughout charging and discharging. Previous research has hypothesized two concurrent effects that may drive Mn dissolution in cells during cycling: acid-induced disproportionation of Jahn–Teller active Mn 3+ and structural rearrangement of the cathode lattice. Through depth probing of the Mn oxidation state in both the cathode and separator via soft X-ray absorption spectroscopy (XAS), hard XAS, and X-ray photoelectron spectroscopy (XPS) in progressive states-of-charge, as well as extended X-ray absorption fine structure (EXAFS) analysis of the local Mn environment, the primary driving force of Mn dissolution is determined to be high-voltage structural rearrangement above 4.2 V. Mn dissolution is, additionally, a main source of capacity fade in Li 4 Mn 2 O 5 DRX cells, which retain only 59% capacity after 20 cycles.

Theibault, Monica↗

Dataset for First Full Dalitz Plot Measurement in Neutron β-Decay using the Nab Spectrometer and Implications for New Physics

The Nab apparatus at the Fundamental Neutron Physics Beamline at the Spallation Neutron Source was designed to measure key observations in neutron beta decay, test the Standard Model's description of the weak interaction, and search for new physics. This data was collected using the Nab apparatus and are presented in the article "First Full Dalitz Plot Measurement in Neutron β-Decay using the Nab Spectrometer and Implications for New Physics." This data publication includes CSV (comma-separated values) files which are used to generate Figures 3 - 11 in the linked journal article.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Measuring Local Turbulence Along the Optical Path: Multi-Beam Optical Seeing Sensor

Deflection of light along the optical path is a major source of image degradation for ground-based telescopes. Methods have been developed to measure upper atmospheric seeing based on models of the turbulence in the atmosphere, but due to boundary conditions, transmission within telescope enclosures is more complex. The Multi-beam Optical Seeing Sensor (MOSS) directly measures the component of the image quality degradation from inhomogeneity of the index of refraction within the telescope dome. MOSS outputs four near-parallel beams of light that travel along the optical path and are imaged by the telescope’s detector, landing like starlight on the telescope’s focal plane. By using a strobed light source, we can ‘freeze’ the instantaneous index variations transverse to the optical path. This system captures both ‘dome’ and ‘mirror’ seeing. Through plotting the standard deviation of differential motion between pairs of beams, MOSS enables characterization of the length scale of turbulence within the dome. The temporal coherence of temperature gradients can be probed with different pulse lengths, and the spatial coherence by comparing pairs at different separations across the aperture of the telescope. Optical path turbulence measurements, alongside other telemetry metrics, will guide thermal and airflow management to optimize image quality. A MOSS prototype was installed in the 1.2[Formula: see text]m Auxiliary Telescope (AuxTel) at the Vera C. Rubin Observatory in Chile, and preliminary data constrain the optical path turbulence with a lower bound of 1.4 arcsec. The optical path turbulence varied throughout the night of observing.

Astronomical seeing↗

Spacetime pq theory for AC and DC electric power systems

The 50/60 Hz alternating current (AC) electric power has been the standard and most flexible energy source powering our modern societies for one and a half centuries since the war of the currents: AC versus direct current (DC). A reactive power concept that was introduced at the beginning of the AC power was very useful for circuit/system analysis, design, control, optimization, and ultimately for more efficient and stable generation, transmission, distribution, and consumption. The initial reactive power theory was based on single-phase sinusoidal AC power to capture inductive and capacitive power that yields to net-zero average power over one fundamental cycle. Soon it was expanded to non-sinusoidal AC power and finally to instantaneous three-phase AC power. However, these reactive power theories remain separate and limited to special cases and have never been consolidated and made valid to all cases. Today, more widespread adoption of power electronics and renewable energy is bringing back DC power into the electric grids. The reactive power concept has never been applied to DC power systems. There is no reactive power in DC power systems according to the existing reactive power theories. Do DC power systems really have no reactive power? Capacitors and inductors are widely used in DC just like in AC power systems. Are they not reactive power components? Why are they different from their AC counterparts? Furthermore, are batteries active or reactive power components? What about active devices like power converters (or inverters) with AC (or DC) on one side and DC (or AC) on the other? Do they generate or consume reactive power? Finally, what about AC and DC hybrid power systems? How to define reactive power in such a complex power system that has a multitude of loads, buses, and sources? Is there reactive power between any two loads, any two buses, or any two sources in a power system and what is the total reactive power in such a complex power system as a whole? As the motivation and goal of this paper to answer the above basic questions, to unify the existing AC reactive power theories and to ultimately provide theoretical and insightful guidance for system analysis, design, control, efficiency, optimization, and operation of complex power systems, a concept of spacetime (both spatial and temporal) active and reactive power (pq) theory—the spatiotemporal aspect of active and reactive power—is developed for both AC and DC power systems. The theoretical definitions and physical meanings of the spacetime reactive power will be developed, and real applications and thought experiments/cases/exercises will be explored and discussed. The developed mathematics to define the active (or real) and reactive (or imaginary) power— p and q respectively by dot (scalar) and cross (vector) products of multi-dimension spacetime vectors and time-space mapping principle/law can have some fundamental implications as well.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array↗

FY25 Mid-Year Report: FNCL Enhancements Implementation

During the first half of FY25 the FNCL team has made consistent progress toward the completion of our project goals. The FNCL prototype panel design has been successfully applied to a fully instrumented 3-panel system which is actively under construction. The FNCL Demonstrator System contains solid scintillators instrumented with SiPMs, which operate on an updated CAEN digitizer, requires no high-voltage, and has a smaller overall footprint. The onboard software will include the LLNL-developed GMM-PSD signal processing. Later this year the system will be experimentally tested alongside the baseline FNCL instrument at LLNLs ISSA facility. In addition to a full systems test, the performance of a DD generator for active interrogation measurements compared to the standard AmLi source will be established for both systems. The data collected at the ISSA facility will be used to experimentally validate the FNCL-Fast Isotopic Fuel Assay’s (FIFA) capability to measure U-235 loading and to predict gadolinium poison content with passive interrogation. The FNCL-FIFA modal was benchmarked with simulation-based data and a user-friendly GUI was added earlier this year. Three separate codes have been submitted to the LLNL ESW system for review prior to their transfers. These include the Predictive Modeling Response toolkit, GMM-PSD firmware beta version, and the FNCL-FIFA analysis package with GUI and user documentation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Autonomous organic synthesis for redox flow batteries via flexible batch Bayesian optimization

Traditional trial-and-error methods for materials discovery are inefficient to meet the urgent demands posed by the rapid progression of climate change. This urgency has driven the increasing interest in integrating robotics and machine learning into materials research to accelerate experimental learning. However, idealized decision-making frameworks to achieve maximum sampling efficiency are not always compatible with high-throughput experimental workflows inside a laboratory. For multi-step chemical processes, differences in hardware capacities can complicate the digital framework by introducing constraints on the maximum number of samples in each step of the experiment, hence causing varying batch sizes in variable selection within the same batch. Therefore, designing flexible sampling algorithms is necessary to accommodate the multi-step synthesis with practical constraints unique to each high-throughput workflow. In this work, we designed and employed three strategies on a high-throughput robotic platform to optimize the sulfonation reaction of redox-active molecules used in flow batteries. Our strategies adapt to the multi-step experimental workflow, where their formulation and heating steps are separate, causing varying batch size requirements. By strategically sampling using clustering and mixed-variable batch Bayesian optimization, we were able to iteratively identify optimal conditions that maximize the yields. Our work presents a flexible approach that allows tailoring the machine learning decision-making to suit the practical constraints in individual high-throughput experimental platforms, followed by performing resource-efficient yield optimization using available open-source Python libraries.

Tamura, Clara [Univ. of Washington, Seattle, WA (U↗

Improving the Capabilities and Computational Efficiency of the RTE+RRTMGP Radiation Code (Final Report)

This report details progress on the RTE+RRTMGP radiation codes made during the period of performance. RTE+RRTMGP is a set of codes for computing radiative fluxes in planetary atmospheres. RRTMGP uses a k-distribution to provide an optical description (absorption and possibly Rayleigh optical depth) of the gaseous atmosphere, along with the relevant source functions, on a pre-determined spectral grid given temperatures, pressures, and gas concentration. RTE computes fluxes given spectrally-resolved optical descriptions and source functions. Spectrally-resolved fluxes are summarized (“reduced”) via a user extensible class. The initial release of the code and the design choices are described in Pincus et al. 2019; the codes are available on Github. Although RRTMGP was based on current (at the time) empirical spectroscopic data, RTE and RRTMGP were developed in large part to modernize software practices. The design focused on flexibility broadly interpreted: by separating code from data and allowing data to drive computation; in coupling to the host model (e.g. the coupling of clouds to radiative fluxes is user-controlled); with respect to programming languages (computational tasks are accessed via widely-compatible C interfaces); and with respect to hardware (the codes run on a range of CPU and GPU architectures). The code also puts an emphasis on modularity and clarity. RTE+RRTMGP v1.0 was released in September 20219. This award supported the evolution of the RTE+RRTMGP code base to support greater flexibility, accuracy, and efficiency.

54 ENVIRONMENTAL SCIENCES↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗