Search NASA⌕ Search

SEARCH · Search NASA

Results for “Kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

ENDF/B-VIII.1

The ENDF/B-VIII.1 release is the newest evaluated nuclear data library produced, distributed, and recommended by CSEWG for use in nuclear science and technology applications. Among the many key advances, relative to the previous version ENDF/B-VIII.0, are: re-evaluation of 239Pu file by a joint international effort; updated 16,18O, 19F, 28-30Si, 50-54Cr, 55Mn, 54,56,57Fe, 63,65Cu, 139La, 233,235,238U, and 240,241Pu neutron nuclear data by the IAEA-coordinated INDEN collaboration; significant changes for 3He, 6Li, 9Be, 51V, 88Sr, 103Rh, 140,142Ce, Dy, 181Ta, Pt, 206-208Pb, and 234,236U neutron data; new nuclear data for the photo-nuclear, being 196 adopted from the IAEA2019 Photonuclear Data Library and one new file from JENDL-5; and new evaluations for the charged-particle and atomic sublibraries. Numerous thermal neutron scattering kernels were re-evaluated or provided for the very first time. Additionally, new covariance testing was implemented. ENDF/B-VIII.1 reduced bias in the simulations of many integral experiments with particular progress noted for fluorine, copper and stainless steel containing benchmarks. Data issues which had hindered the deployment of ENDF/B-VIII.0 for commercial nuclear power applications in high burn-up situations, were addressed. ENDF/B-VIII.1 data are distributed in both ENDF-6 and GNDS formats.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Tropical High Cloud Feedback Relationships to Climate Sensitivity

Clouds constitute a large portion of uncertainty in predictions of equilibrium climate sensitivity (ECS). While low cloud feedbacks have been the focus of intermodel studies due to their high variability among global climate models, tropical high cloud feedbacks also exhibit considerable uncertainty. Here, we apply the cloud radiative kernel technique of Zelinka et al. to 22 models across the CMIP5 and CMIP6 ensembles to survey tropical high cloud feedbacks and analyze their relationship to ECS. We find that the net high cloud feedback and its altitude and optical depth feedback components are significantly positively correlated with ECS in the tropical mean. On the other hand, the tropical mean high cloud amount feedback is not correlated with ECS. These relationships are most pronounced outside of areas of strong climatological ascent, suggesting the importance of thin cirrus feedbacks. Finally, we explore connections between high cloud feedbacks, climate sensitivity, and mean state high cloud properties. In general, high ECS models are cloudier in the upper troposphere but have a thinner high cloud population. Furthermore, we find that having more thin cirrus in the mean state relates to more positive high cloud altitude and optical depth feedbacks, and it either amplifies or dampens the high cloud amount feedback depending on the large-scale dynamical regime (amplifying in descent and dampening in ascent). In summary, our analysis highlights the importance of tropical high cloud feedbacks for driving intermodel spread in ECS and suggests that mean state high cloud characteristics might provide a unique opportunity for observationally constraining high cloud feedbacks.

atmosphere↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Disparities in the air quality monitoring stations and PM₂.₅ in Chicago’s air quality landscape

Fine particulate matter (PM₂.₅) poses significant public and environmental health risks in urban areas. Chicago’s dense industry and traffic create variable air quality, yet monitoring is unevenly distributed, resulting in undersampling of air quality data in some city areas. This study applied a hybrid approach using GIS-based kernel density mapping, interpolation modeling (IDW, Spline, Kriging) of USEPA monitoring data, multi-scale temporal trend analyses (hourly to annual), and ESDA. Accordingly, the density surface showed that monitors are concentrated in the affluent north, northwest, and southwest sides of Chicago (up to ~ 0.07 stations per sq mile), while the south and southeast regions, with predominantly minority communities, have virtually no coverage. Overall, citywide coverage is minimal (~ 4–5 monitors total; ~0.02 per sq mile; ≈1 per 600,000 residents). Temporal analyses showed that the city’s mean annual PM₂.₅ (~ 10.8 µg/m³) exceeds USEPA/WHO standards (9 µg/m³), with summer means (~ 17.1 µg/m³) significantly higher than other seasons. Diurnally, a clear pattern was observed, with PM₂.₅ concentrations peaking overnight (00:00–03:00) and during the morning rush hours, and dipping during midday to late afternoon. Spatial distribution of PM₂.₅ identified hotspots near O’Hare Airport, the downtown Loop area, and south-side neighborhoods, contrasting with lower concentrations on the north side, revealing Chicago’s socioeconomic divides and resulting environmental inequities. The findings underscore the need for expanded monitoring and targeted interventions in under-monitored, high-pollution communities to advance equitable community health.

54 ENVIRONMENTAL SCIENCES↗

Micromechanical Properties of the SiC and Pyrolytic Carbon Layers in Tristructural-Isotropic Coated Particles

Tristructural isotropic (TRISO) coated particle fuel was initially developed for high temperature gas cooled reactors (HTGR) and has been proposed for several other advanced reactor concepts. The design of TRISO particles focuses on preventing the release of fission products in normal and off-normal reactor conditions. The particle design features an actinide bearing fuel kernel that is surrounded by three pyrolytic carbon (PyC) layers and a silicon carbide layer (SiC). The mechanical stability of the particle and fission product retention for both metallic and gaseous fission products depend on the SiC layer. Post irradiation examination (PIE) of TRISO fuel from the first two US DOE Advanced Gas Reactor Fuel Development and Qualification Program irradiation campaigns, AGR-1 and AGR-2, had identified a low rate of particles exhibiting cracking in the SiC layer that did not propagate across the SiC layer. While cracking in the SiC is rare for test conditions and particles associated with the AGR program, understanding the stress state and mechanical properties of the SiC and PyC layers related to particle architecture can aid predicting thermomechanical response of TRISO fuel under the prescribed operation envelope and beyond as well as aiding in the development of similar fuel concepts for other advanced reactors. The presented investigation shows the relationship of the mechanical properties and mechanical response (e.g., understanding crack propagation) of the SiC and PyC layers relative to position within the particle. Testing was conducted on the inner and outer PyC layers of TRISO particles to quantify differences in mechanical behavior.

Montoya, Katherine [ORNL] (ORCID:0000000326955086)↗

Nonlinear analog processing with anisotropic nonlinear films

Digital signal processing is the cornerstone of several modern-day technologies, yet in multiple applications it faces critical bottlenecks related to memory and speed constraints. Thanks to recent advances in metasurface design and fabrication, light-based analog computing has emerged as a viable option to partially replace or augment digital approaches. Several light-based analog computing functionalities have been demonstrated using patterned flat optical elements, with great opportunities for integration in compact nanophotonic systems. So far, however, the available operations have been restricted to the linear regime, limiting the impact of this technology to a compactification of Fourier optics systems. In this paper, we introduce nonlinear operations to the field of metasurface-based analog optical processing, demonstrating that nonlinear optical phenomena, combined with nonlocality in flat optics, can be leveraged to synthesize kernels beyond linear Fourier optics, paving the way to a broad range of new opportunities. As a practical demonstration, we report the experimental synthesis of a class of nonlinear operations that can be used to realize broadband, polarization-selective analog-domain edge detection.

analog image processing↗

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES↗

Ristra Project FY23 L2 Milestone Report, Rev.1: MRT #8541: Multiphysics Scaling on EAS-3

The findings of this report were used to close out the ATDM milestone MRT# 8541, which was designed to demonstrate readiness of ATDM multiphysics codes for mission-relevant work on ATS-4, El Capitan. To this end, the closure criteria were to run a 3D shaped charge problem at scale up to 50% of the El Capitan early-access system, RZVernal (AMD Trento CPUs and AMD MI-250X GPUs), demonstrate scalability, and document challenges with the software stack and environment. LANL’s approach to this milestone was to test our modular software capability by developing an entirely new code, Moya, built upon our FleCSI framework. The physics capability and the GPU infrastructure needed for the shaped charge problem on GPUs was added to Moya, and the required calculations were performed at scale. Moya showed good scaling without any fine-tuning of GPU kernels; there is still significant room for performance enhancements, especially for the Legion backend. Tied up in this L2 milestone was a closeout of KPP-3s for the ECP ST Projects at LANL; this material will be covered in a separate document.

97 MATHEMATICS AND COMPUTING↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Blueprints for Training Information Bottlenecks for Collider Analyses

Dimensionality reduction is a crucial aspect of data analysis in high energy physics, even if accompanied by information loss. Several methods, including histogram- and kernel-based analyses, are only computationally feasible for low-dimensional data. Furthermore, simulation models used in HEP can often only be validated for low-dimensional data. We provide several blueprints for using machine learning to create low-dimensional data representations (continuous event variables and discrete classification labels) for use in signal discovery and parameter estimation tasks. We also describe how to design the learned representation to facilitate a) searches with unknown model parameters and b) validation of simulation models in data control regions.

43 PARTICLE ACCELERATORS↗

PV-Finder: ML Based Algorithm for Primary Vertex Identification

he CMS detector at the High-Luminosity Large Hadron Collider (HL-LHC) will operate in challenging conditions with expected pile-up of up to 200 collisions per bunch crossing, necessitating the development of a more resilient primary vertex (PV) reconstruction method to ensure the integrity of data analysis and the efficiency of the CMS triggering system. This contribution describes preliminary studies on a new ML based PV-Finder method for PV identification. The method is based on a model trained using Kernel Density Estimations (KDEs) derived from the positions of reconstructed tracks at the beamline, incorporating uncertainties from track parameters. It also utilizes target histograms, modeled as Gaussian distributions centered on the actual ground truth values of specific primary vertices.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Evaluation of Radiography for TRISO Buffer Layer Density Measurement

Tristructural isotropic (TRISO) fuel particles consist of a central uranium-bearing kernel and a series of coating layers designed to retain fission products and to ensure fuel performance. Several parameters such as thickness and density must be measured for these coating layers to show that they conform with fuel specifications. Current methods for measuring the density of pyrolytic carbon and silicon carbide layers (liquid gradient density column) and the buffer layer (mercury porosimetry) generate Resource Conservation and Recovery Act (RCRA) radiological-mixed waste. In addition, measurement of buffer and inner pyrolytic carbon layer densities require hot sampling or interrupted coating runs and the mercury porosimetry method used for buffer density measurement only measures the mean buffer density, not the interparticle distribution. A new approach has been evaluated to measure the density of coating layers in TRISO particles based on the dependence of x-ray attenuation in radiographs on material density. This method does not generate RCRA mixed waste, measures density on a particle-by-particle basis, and in principle is capable of measuring the density of all coating layers in a single process. Initial results using thinned TRISO particle sections to evaluate radiography measurement of density as a quality control characterization method are reported herein. In this work, the primary focus is on measurement of the density of the buffer layer; however, with appropriate calibration the method should be applicable to other coating layers. Improvements to the initial method and a full demonstration of the method on the remaining coating layers may be pursued as a future effort.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Evaluation of XCT for Matrix Density Measurement of Particle Fuel Forms

Particle fuel forms generally consist of a dispersion of fuel, such as tristructural isotropic (TRISO) particles, within a refractory matrix (e.g., graphite or silicon carbide). The density of matrix materials for particle fuel forms is of interest for modeling fuel form strength and thermal properties and may be specified as a quality control parameter, depending on reactor design. Some of the uncertainty associated with traditional, manual approaches can be eliminated by performing x-ray computed tomography (XCT) on the fuel forms and applying image processing methods to generate a precise count of the number of particles. This also removes the need to include determination of particle count within each individual fuel form during fabrication. Unfortunately, reconstruction artifacts from high-Z uranium-bearing kernels prevent accurate measurement of individual particle volumes using this approach, so the use of mean particle mass and volume are still necessary for computation of average fuel form matrix density. This method of using XCT to count particles in individual fuel form for determination of average matrix density was applied to three archived compacts from the Advanced Gas Reactor Fuel Development and Qualification (AGR)-1 campaign, four archived compacts with uranium carbide/uranium oxide (UCO) TRISO from the AGR-2 campaign, and three archived UO 2 -TRISO compacts from the AGR-2 campaign. The resulting density values were compared with those previously reported, showing slight changes due to uncertainties in the previously used number of particles in each of these cylindrical, graphite matrix compacts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Performance Results on CPU/GPU Exascale Architectures for OMEGA: The Ocean Model for E3SM Global Applications

The US Department of Energy (DOE) conducts climate simulations on some of the world’s largest supercomputers. These exascale machines use heterogeneous architectures with both CPUs and GPUs, and scientific codes must adapt to make full use of this computing power. Los Alamos National Lab is developing Omega: The Ocean Model for E3SM Global Applications, which is specifically designed for modern exascale computers. It uses external libraries that have been optimized for a variety of architectures to run on different supercomputers. Omega is an unstructured-mesh ocean model based on TRiSK numerical methods. It will be the new ocean component of the DOE’s Energy Exascale Earth System Model (E3SM). The algorithms in Omega follow those of the current ocean component, MPAS-Ocean, but it will be written in C++ rather than Fortran to take advantage of the Kokkos performance portability library. Omega spatial operators are written as Kokkos kernels to run efficiently on both CPUs and GPUs. Work on Omega began in 2023 with a new C++ framework for unstructured mesh partitioning, halo exchanges, parallel IO, and Kokkos interfaces. The current version, Omega-0, is being developed to solve the shallow water equations and at present includes all of the tendency terms but not time stepping. Here we share the results of Omega-0 verification and performance testing. Verification includes unit tests implemented with CTest as well as convergence tests in Polaris, an in-house python package with a large suite of test problems. Performance tests compare simulations conducted on CPUs versus GPUs and across different architectures: tests are run on Frontier, which has AMD “Optimized 3rd Gen EPYC” CPUs and AMD MI250X GPUs, as well as Perlmutter, which is composed of AMD EPYC 7763 CPUs and NVIDIA A100 GPUs.

58 GEOSCIENCES↗