Search NASASearch

SEARCH · Search NASA

Results for “Mesh optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

High order interpolation of magnetic fields with vector potential reconstruction for particle simulations

We propose a method for interpolating divergence-free continuous magnetic fields via vector potential reconstruction using Hermite interpolation, which ensures high-order continuity for applications requiring adaptive, high-order ordinary differential equation (ODE) integrators, such as the Dormand-Prince method. The method provides C(m) continuity and achieves high-order accuracy, making it particularly suited for particle trajectory integration and Poincaré section analysis under optimal integration order and timestep adjustments. Through numerical experiments, we demonstrate that the Hermite interpolation method preserves volume and continuity, which are critical for conserving toroidal canonical momentum and magnetic moment in guiding center simulations, especially over long-term trajectory integration. Furthermore, we analyze the impact of insufficient derivative continuity on Runge-Kutta schemes and show how it degrades accuracy at low error tolerances, introducing discontinuity-induced truncation errors. Lastly, we demonstrate performant Poincaré section analysis in two relevant settings of field data collocated from finite element meshes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Elucidating metal–organic framework structures using synchrotron serial crystallography

Metal organic frameworks (MOFs) are porous crystalline materials that display a wide variety of physical and chemical properties. Their single crystal structure determination is often challenging because in most cases micro- or nano-sized crystals spontaneously form upon MOF synthesis, which cannot be recrystallized. The production of larger single crystals for structure determination involves optimizing, and thus modifying, the conditions of synthesis, in which success cannot be guaranteed. Failure to produce crystals suitable for single-crystal X-ray diffraction leaves the 3D structure of the MOF compound unknown, and scientists must resort to more challenging structure solution methods based on X-ray powder or electron diffraction data. These laborious tasks can be avoided by using serial crystallography techniques which merge data collected on many micro-crystals. Here, we report the application of three synchrotron serial crystallography methods. We call these “mesh”, “grid” and “mesh&collect” scans. “Still” images (no rotation) are collected in the mesh scan approach, whereas small rotational wedges are collected in the grid scan method. The third protocol, mesh&collect, combines the acquisition of still images and rotational wedges. Using these means, we determine the ab initio structure of benchmark MOFs, MIL-100(Fe) and ZIF-8, that differ largely in unit cell size. These methods are expected to be widely applicable and facilitate structure determination of many MOF microcrystalline systems.

36 MATERIALS SCIENCE

ASU’s DAC polymer-enhanced cyanobacterial bioproductivity (AUDACity)

ASU’s DAC polymer-enhanced cyanobacterial bioproductivity (AUDACity) project aims to demonstrate a novel, scalable method for removing carbon dioxide (CO 2 ) directly from ambient air and delivering it to cyanobacterial cultures to produce commodity biofuel, mid-value protein for supplements, and high value phycocyanin (PC), a natural blue colorant (Figure A). This approach uses low-cost, reusable anion exchange polymers embedded in modular mesh packets, which capture CO 2 during drying cycles when exposed to ambient air, and release concentrated CO 2 into aqueous cultivation systems. The project addresses a critical challenge in energy research needed for developing sustainable, economically viable methods of Direct Air Capture (DAC) that can be integrated with bio-based systems for fuel and chemical production. AUDACity contributes to scientific understanding by integrating materials chemistry, cyanobacterial biology, and system engineering to create a distributed CO 2 delivery platform. Key insights have emerged around the design of biocompatible sorbents, optimization of CO 2 capture-release cycles, and durability of packet-based delivery systems under outdoor conditions. Notably, the team has synthesized and tested a range of polymer sorbents, identified mechanisms of material degradation and fouling, and advanced both lab- and pilot-scale cultivation systems to evaluate performance. From a technical and economic standpoint, AUDACity shows promise for achieving cost-effective CO 2 capture and delivery into aqueous media and biofuel production. Preliminary techno-economic analysis (TEA) indicates that the DAC system based on current performance can reach $\$$680/tonne CO 2 delivered into aqueous solution; with reasonable improvements to sorbent lifetime, sorbent capacity, reducing water uptake the approach could reach $\$$66/tonne by avoiding the need for energy-intensive sorbent regeneration and CO 2 compression, making it more feasible for decentralized deployment. With these costs for CO 2 and by extracting and selling high-value PC ($\$$50/kg) and mid-value protein supplement ($\$$6/kg), the remaining biomass can be hydrothermally treated into biofuel for $\$$2.50/gallon, and would support a small first-of-a-kind biorefinery capable of producing 500 barrels per day of biofuel. The project offers meaningful public benefits by advancing carbon removal technologies that are low-energy, modular, and adaptable to non-arable land and brackish water use. It aligns with national goals to develop advanced biotechnology and supports future pathways for bio-based fuels and products. By enabling direct coupling of CO 2 transfer into aqueous medium and biological carbon utilization, AUDACity lays the groundwork for effective algae cultivation without wasteful CO 2 delivery and is a promising and innovative solution for low-carbon fuel and bioproduct generation contributing to a vigorous bioeconomy.

09 BIOMASS FUELS

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian

Numerical Simulation of a Natural Convection–Driven Air-Cooled Reactor Cavity Cooling System Experiment

Ensuring the efficient removal of decay heat from the reactor vessel is essential for the safety of advanced reactor technologies. Several Generation-IV concepts incorporate variations in the reactor vessel cooling systems to achieve this objective. High-temperature gas-cooled reactors utilize a reactor cavity cooling system (RCCS), a passive ex-vessel system designed to operate without active components or external power during accident conditions. The RCCS removes decay heat primarily through radiative and convective heat transfer mechanisms. Here, this study presents a comprehensive validation of a computational fluid dynamics Reynolds-averaged Navier-Stokes model for the University of Wisconsin-Madison air-cooled RCCS facility. Validation was conducted for both high- and low-power natural convection cases under a uniform heating profile. Near-wall resolution was found to be critical for accurately modeling natural convection in the RCCS; employing an all-𝑦 + wall treatment resulted in wall temperature discrepancies exceeding 50 °⁢𝐶 compared to a wall-resolved mesh. Thermal-hydraulic behaviors under natural and forced convection conditions were compared within the heated cavity and RCCS. A turbulence model sensitivity analysis indicated that low-Reynolds number k-ɛ, k-ω shear stress transport (SST), and Reynolds stress transport models produce similar wall temperature predictions. A buoyancy modeling sensitivity study revealed that the Boussinesq approximation significantly underpredicted thermal-hydraulic behavior in the RCCS. Based on these findings, modeling recommendations are provided. The validated data set along with identified sensitivities refine the modeling of natural convection in the RCCS. The information produced by this study supports RCCS design, optimization, and safety evaluations, enabling the calibration and verification of reduced-order thermal-hydraulic models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

AMReX and pyAMReX: Looking beyond the exascale computing project

AMReX is a software framework for the development of block-structured mesh applications with adaptive mesh refinement (AMR). AMReX was initially developed and supported by the AMReX Co-Design Center as part of the U.S. DOE Exascale Computing Project (ECP), and is continuing to grow post-ECP. In addition to adding new functionality and performance improvements to the core AMReX framework, we have also developed a Python binding, pyAMReX, that provides a bridge between AMReX-based application codes and the data science ecosystem. pyAMReX provides zero-copy application GPU data access for AI/ML, in situ analysis and application coupling, and enables rapid, massively parallel prototyping. In this paper we review the overall functionality of AMReX and pyAMReX, focusing on new developments, new functionality, and optimizations of key operations. We also summarize capabilities of ECP projects that used AMReX and provide an overview of new, non-ECP applications.

Myers, Andrew

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce

Optimizing Simulation Fidelity in Direct-Drive Inertial Confinement Fusion with Cassio

Recently, the National Ignition Facility (NIF) demonstrated that inertial confinement fusion (ICF) is capable to achieve thermonuclear (TN) ignition in the laboratory, making it a crucial method on the path to replicate the Sun’s power production mechanism on Earth. However, the physics governing the high-energy density environments is very complex and remains a challenge to fully understand and model. For example, dopants in the TN fuel are important diagnostic tools to extract the thermodynamic conditions of the plasma. However, if their concentration is chosen too high, they can significantly degrade the performance of an ICF capsule. In this study, we use the Los Alamos National Laboratory radiation-hydrodynamics code Cassio to model ICF implosions of capsules which contain deuterium fuel with high-Z dopants like Krypton and Argon from the high-Z campaign conducted 15 years ago. We focus on how chosen computational and physics parameters influence the implosion outcomes. By systematically changing the resolution of the computational mesh and the photon energies as well as modifying settings for the laser drive and TN fuel pre-heat effects, we assess the impact on experimentally measured performance metrics like neutron production from TN burn and x-ray emission during the implosion. Our results will help to improve the fidelity of simulations and guide future numerical studies and experimental designs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Actinium–DOTA coordination in water from hybrid ML/MM: Structure, free energies, and water-exchange pathways

Quantitative simulation of trivalent ƒ-block chelates in water remains challenging because bonded and non-bonded force-field models make different approximations for coordination structure, exchange dynamics, and ion–ligand interactions in highly charged systems. Here, we develop a hybrid machine-learning/molecular-mechanics (ML/MM) framework for Ac 3+ –DOTA in explicit solvent by training an E(3)-equivariant neural network potential (MACELES) on mechanically embedded QM/MM data for Ac aquo and Ac–DOTA species and coupling it to NAMD 2.14 with particle-mesh Ewald electrostatics. Nanosecond ML/MM trajectories remain numerically stable and preserve chelate integrity, yielding a compact DOTA inner shell with an inner-sphere water coordination number of CN Ac,O w ≈ 1.7 arising from a dynamic equilibrium between one- and two-water states (37.5% and 59.9% of frames; three waters 2.5%). A 5 ns potential of mean force shows two low-lying basins at CN Ac,O w ≈ 1 and CN Ac,O w ≈ 2. DFT end-state free energies are consistent with the ML/MM profile, and DFT minimum-energy paths provide a qualitative electronic-structure reference for the observed basin connectivity. State-resolved kinetics reveal picosecond water-exchange pathways that couple hydration changes to transient DOTA arm fluctuations, and training-set comparisons show that temperature-matched Ac–DOTA data optimize energy/force accuracy while more diverse solvated data improve charge prediction. Overall, the present hybrid ML/MM model provides a practical description of Ac 3+ –DOTA hydration thermodynamics and short-time exchange behavior in explicit water at MD-like cost.

Actinium

Integrated Neutronics Modeling for Inertial Fusion Energy Systems: Development and Application to LD-FIRST

Lawrence Livermore National Laboratory (LLNL) is proposing a new Laser Driven Fusion Integration Research and Science Test Facility (LD-FIRST) with the goal of providing an experimental testbed for future Inertial Fusion Energy (IFE) systems. However, IFE systems require detailed and accurate multiphysics modeling to quantify material damage, thermal loading, and tritium breeding within complex chamber environments. This article presents the first step in an integrated multiphysics framework that couples meshed CAD-based geometry within Monte Carlo neutronic simulations to enable high-fidelity analysis of IFE chamber concepts, with future coupling to external codes. The neutronics workflow utilizes OpenMC and its third-party capability to use CAD-based geometries through DAGMC and tally on unstructured meshes with Libmesh to evaluate neutron transport behavior, geometric fidelity, and material performance under reactor-relevant conditions. The use of tailored tallies on unstructured meshes in this framework allows direct transfer without interpolating to CFD simulation tools. Two IFE chambers were evaluated, both conceived by LLNL: HYLIFE-II and Laser IFE (LIFE). This work produced high-fidelity conformal surface and volumetric meshes of the HYLIFE-II and LIFE chambers with mapped spatial insight into material damage, thermal loading, and tritium breeding. The HYLIFE-II model was built utilizing available resources and used as a test case to verify that the neutronics framework can handle complex geometries. The LIFE chamber CAD was provided by LLNL and was the main focus of this work. This work analyzes multiple ternary alloy breeding materials for the LIFE chamber, across different 6 Li enrichments to produce data relevant to the LD-FIRST project. This work also investigates the level of model fidelity for the LIFE chamber, and results show that inclusion of detailed first wall and coolant structures increased the predicted tritium breeding ratio (TBR) by ~30%, highlighting the sensitivity of tritium breeding and the need for a high-fidelity simulation framework for IFE chambers. These developments provide a scalable toolset for the design and optimization of next-generation IFE chambers, forming a solid foundation for future coupled multiphysics analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Reduced-Order CFD Modeling to Support Waste Loading Optimization in Hanford WTP Vitrification

The U.S. DOE Hanford Site stores over 56 million gallons of radioactive liquid tank waste that must be treated and immobilized for long-term disposal The Waste Treatment and Immobilization Plant (WTP) will vitrify this waste by feeding it into Joule-heated melters, where it is incorporated into a stable borosilicate glass Computational fluid dynamics (CFD) simulations of glass melters can provide insight into the maximum achievable waste loading under varying melter operating conditions Fully resolved VOF multiphase simulations were used as the reference model to capture bubble-driven convection in the melter, including bubble formation, rise behavior, and induced glass melt circulation Effective bubble column diameter and rise velocity were extracted from the resolved simulations, compared with empirical correlations, and refit across relevant viscosity and gas flow rate conditions Explicit gas–liquid interface tracking was replaced with a single-phase momentum source term model, enabling faster steady-state CFD simulations while preserving the dominant hydrodynamic effects of bubbling New empirical correlations were developed for effective bubble column diameter and bubble rise velocity by fitting resolved simulation data across expected melter viscosity and gas flow rate ranges, providing improved inputs for the momentum source term model compared with existing literature correlations The momentum source term model reduced fluid-domain mesh size by 89% and achieved an 8.4× computational speedup relative to resolved bubbling simulations The validated momentum source term approach enables prediction of process-relevant heat transfer behavior in the integrated melter model, including heat transfer from the molten glass to the cold cap, plenum, refractory walls, and surrounding structural regions under varying melter operating conditions

12 - MGMT OF RADIOACTIVE AND NON-RADIOACTIVE WASTE

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING

Openpronghorn

OpenPronghorn is a simulation tool specifically tailored for modeling thermal-hydraulic phenomena in advanced nuclear reactors. It is built on the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source platform that facilitates the development of high-performance scientific computing applications. OpenPronghorn solves the Navier-Stokes equations, which describe the conservation of mass, momentum, and energy in fluid flows, using the finite volume numerical method. The code supports a wide range of fluid flow conditions that are applicable to nuclear reactors, including incompressible and weakly compressible flows, as well as single-phase and multiphase flows. It is capable of modeling diverse flow regimes, including laminar and turbulent flows, using various turbulence models such as the standard k-epsilon models, the v2f model, and the mixing length model. For multiphase flows, OpenPronghorn employs a mixture a Eulerian modeling approach with mixture, drift-flux, and full Eulerian models, and includes open-sourced interfacial transfer correlations for drag, exchange, and heat transfer coming from the scientific literature. OpenPronghorn's modular design allows it to handle multiscale simulations, ranging from detailed Reynolds-Averaged Navier Stokes (RANS) simulations to coarse-mesh and lumped parameter models. This flexibility enables users to perform high-fidelity simulations of specific reactor components as well as system-level analyses of entire reactor circuits. The code can be coupled with other MOOSE-based tools using the MultiApp system, allowing for the transfer of coupling quantities such as mass flow rates, heat fluxes, and boundary conditions between different simulation scales. One of the main features of OpenPronghorn is the it includes built-in validation cases from the open-source scientific literature and supports the implementation of user-defined models and correlations through MOOSE's FunctorMaterial system. OpenPronghorn is designed to be computationally efficient, leveraging the SIMPLE projection method for large-scale problems, and can be run on high-performance computing systems to handle the extensive computational demands of detailed reactor simulations. Overall, OpenPronghorn is a versatile and robust tool that provides critical insights into the thermal-hydraulic behavior of advanced nuclear reactors, supporting the design, safety, and optimization of next-generation nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L

Towards a NEAMS-based high-fidelity model of the MARVEL reactor

This report outlines the progress of Idaho National Laboratory in developing a high-fidelity and high-resolution model of the Microreactor Applications Research Validation and Evaluation reactor. The model was developed under the Nuclear Energy Advanced Modeling and Simulation microreactor application driver at Idaho National Laboratory. The overarching objective of this activity is the development of a high-fidelity multiphysics MARVEL model using NEAMS tools, and to verify and validate NEAMS tools against MARVEL reference simulation and experimental data, respectively. This is a unique opportunity to conduct multiphysics analysis on a soon-to-be-deployed microreactor. This multiphysics model developed under the NEAMS-funded INL microreactor application driver leverages three single-physics models coupled via the MOOSE’s MultiApp and Transfer systems. The latter systems enable in-memory data transfer between MOOSE-based and MOOSE-wrapped applications. The first single-physics model, that functions as main application, leverages Griffin to model the neutron transport in the core through the discontinuous finite element (DFEM) discrete ordinates solver (SN). Several optimization flags that were developed by the Griffin developer team were beta-tested to enhance the solver’s performance. These include the combined use of using_average_xs and update_averaged_xs_on that enable to avoid expensive on-the-fly cross sections evaluations at each linear iterations in favor of evaluations of the macroscopic cross sections at each Picard iteration. The second single-physics model uses BISON to handle solid heat transfer and asymptotic hydrogen redistribution analysis in the fuel. While the model returns consistent results for the temperature and hydrogen distribution in the fuel, a mismatch was noticed in the calculated temperature in the reflector due to the value of the gap conductance used in our model. Ongoing investigations are being performed to assess the origin of this discrepancy. Finally, the System Analysis Module (SAM) was used to model the flow of the sodium-potassium eutectic in the primary loop. A first verification was also performed showing good agreement in terms of mass flow rate and inlet temperature. All mesh files were generated using the MOOSE Reactor module, removing the need for external meshing tools. Notably, this workscope represents one of the initial applications of the MOOSE Reactor module for modeling highly irregular geometries. The use of the reactor module significantly streamlined the mesh generation process. The full multiphysics mode, that combines all the single physics models, was leveraged to conduct initial steady-state multiphysics simulations to compute power, and temperature distribution in the reactor. Initial testing was performed for transient simulations as well. In this case, the new checkpoint restart capability for eigenvalue calculations was tested showing the capability for streamlined restart of transient calculations. Future work will focus on improving the fidelity of the model by performing comprehensive code-to-code comparisons. For instance, the full-core Griffin neutronics model will be benchmarked against MCNP reference results, that were provided by the MARVEL design team. Additionally, the SAM T/H model will be verified against reference RELAP-5 results for selected accident scenarios. Besides code-to-code verification exercises, the model fidelity will be improved by replacing the single-channel SAM model with a more complex SAM-Pronghorn coupled model, in which the sub-channel capability is deployed to obtain radial temperature resolution in the coolant. This model will be developed in synergy with the NEAMS thermal hydraulics team.

22 GENERAL STUDIES OF NUCLEAR REACTORS

DS-GL: Advancing Graph Learning via Harnessing the Power of Nature within Dynamic Systems

With the rapid digitization of the world, an increasing number of real-world applications are turning to nonEuclidean data, modeled as graphs. Due to their intrinsic high complexity and irregularity, learning from graph data demands tremendous computational power. Recently, CMOS-compatible Ising machines, i.e., dynamic systems composed of CMOS components, have emerged as a new approach that harnesses the inherent power of natural annealing within dynamic systems to efficiently resolve binary optimization problems and have been adopted for traditional graph computation, such as max-cut. However, when performing complex Graph Learning (GL) tasks, Ising machines face significant hurdles: (i) they are inherently binary and thus ill-suited for real-valued problems; (ii) their expensive all-to-all coupling network that guarantees effective natural annealing poses daunting scalability concerns. To address these challenges, this paper proposes a nature-powered graph learning framework dubbed DS-GL, which is the first effort to transform the process of solving graph learning problems into the natural annealing process within a parameterized dynamic system embodied as a CMOS chip. To tackle the two major hurdles, DS-GL first augments the Ising machine architecture to modify the self-reaction term of its Hamiltonian function from linear to quadratic, effectively serving as an energy regulator. This adjustment maintains the system’s original physical interpretation while enabling it to process continuous, real-valued data. Second, to address the scaling issue, DS-GL further upgrades the real-valued dense Ising machine by decomposing it into a mesh-based multi-PE dynamic system that supports efficient distributed spatial-temporal co-annealing across different PEs through sparse interconnects. By exploiting the inherent sparsity and component structures in real-world graphs, DS-GL is able to map complex graph learning tasks onto the scalable dynamic system while maintaining high accuracy. Evaluations with three diverse GL applications across six real-world datasets, including traffic flow and COVID-19 prediction, show that DS-GL can deliver from 102× to 106× speedups and 500× energy reduction over Graph Neural Networks on GPUs, with 5% - 20% accuracy enhancement.

Song, Ruibing

Bias-Variance Trade-Off in Physics-Informed Neural Networks with Randomized Smoothing for High-Dimensional PDEs

Physics-Informed Neural Networks (PINNs) have triggered a paradigm shift in scientific computing, leveraging mesh-free properties and robust approximation capabilities. While proving effective for low-dimensional partial differential equations (PDEs), the computational cost of PINNs remains a hurdle in high-dimensional scenarios. This is particularly pronounced when computing high-order and high-dimensional derivatives in the physics-informed loss. Randomized Smoothing PINN (RS-PINN) introduces Gaussian noise for stochastic smoothing of the original neural net model, enabling the use of Monte Carlo methods for derivative approximation, which eliminates the need for costly automatic differentiation. Despite its computational efficiency, especially in the approximation of high-dimensional derivatives, RS-PINN introduces biases in both loss and gradients, negatively impacting convergence, especially when coupled with stochastic gradient descent (SGD) algorithms. We present a comprehensive analysis of biases in RS-PINN, attributing them to the nonlinearity of the Mean Squared Error (MSE) loss as well as the intrinsic nonlinearity of the PDE itself. We propose tailored bias correction techniques, delineating their application based on the order of PDE nonlinearity. The derivation of an unbiased RS-PINN allows for a detailed examination of its advantages and disadvantages compared to the biased version. Specifically, the biased version has a lower variance and runs faster than the unbiased version, but it is less accurate due to the bias. To optimize the bias-variance trade-off, we combine the two approaches in a hybrid method that balances the rapid convergence of the biased version with the high accuracy of the unbiased version. In addition to methodological contributions, we present an enhanced implementation of RS-PINN. Extensive experiments on diverse high-dimensional PDEs, including Fokker-Planck, Hamilton-Jacobi-Bellman (HJB), viscous Burgers’, Allen-Cahn, and Sine-Gordon equations, illustrate the bias-variance trade-off and highlight the effectiveness of the hybrid RS-PINN. Empirical guidelines are provided for selecting biased, unbiased, or hybrid versions, depending on the dimensionality and nonlinearity of the specific PDE problem.

97 MATHEMATICS AND COMPUTING

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability