Search NASA⌕ Search

SEARCH · Search NASA

Results for “High Performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

Density functional theory-based surrogate kinetic models for heterogeneous reactions of hydrocarbon intermediates on silicon carbide

The increasing demand for high-performance materials in advanced technologies highlights the importance of achieving a fundamental understanding and potential control of silicon carbide (SiC) deposition processes. However, existing models often lack sufficient theoretical detail, relying heavily on empirical data and offering limited predictive capability. In particular, the complex surface chemistry governing SiC growth remains poorly understood. This study addresses these challenges by employing density functional theory (DFT) to investigate key heterogeneous reactions involving hydrocarbon intermediates on SiC surfaces, including dehydrogenation, hydrogenation, and carbon deposition. Transition state searches were conducted to identify reaction pathways and energy barriers. While first-principles calculations offer high accuracy, they are computationally intensive. To extend the utility of these first-principles results, vibrational analyses were performed using phonon-based statistical thermochemistry to compute temperature-dependent reaction rates which were used to develop Arrhenius-type surrogate kinetic models. Furthermore, the resulting framework provides a more rigorous, physically grounded basis for integrating atomistic insights into continuum-scale modeling, ultimately enabling improved prediction and optimization of SiC film growth in high-performance material systems.

Density Functional Theory↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

International fuel performance study of fresh fuel experiments for PCMI effects during RIA experiments

This paper presents the results of High-burnup Experiments for Reactivity-initiated Accident (HERA) Modeling & Simulation (M&S) exercise. The HERA project under the Nuclear Energy Agency (NEA) Second Framework for Irradiation Experiments (FIDES-II) program is focused on studying Light Water Reactor (LWR) fuel behavior during Reactivity-Initiated Accident (RIA) conditions. The Part I M&S cases are based on a series of tests in the Transient Reactor Test (TREAT) facility in the United States and the Nuclear Safety Research Reactor (NSRR) in Japan. The purpose of this work is to evaluate the test design to accomplish its goals in establishing clearer understanding of the effects of power pulse width during RIA conditions. Further, the blind predictions using various computational tools have been performed and compared amongst to interpret the behaviors of high burnup fuels during RIA. While many international participants evaluate the thermal–mechanical behavior of fuel rod under different conditions, a considerable scatter of outputs comes out for the cases due to the disparity between codes in predicting mechanical behaviors. In general, however, the results of thermal–mechanical analysis elaborate that nominal design conditions the shorter pulse width tests in NSRR should cause cladding failures while the TREAT tests appear to have more split prediction of failure or not. Furthermore, the sensitivity analysis varying key testing parameters reveals the considerable effect of power pulse width and total energy deposition on prediction of fuel rod failure.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Remark on Algorithm 1012: Computing Projections with Large Datasets

In ACM TOMS Algorithm 1012, the DELAUNAYSPARSE software is given for performing Delaunay interpolation in medium to high dimensions. When extrapolating outside the convex hull of the training set, DELAUNAYSPARSE calls the nonnegative least squares solver DWNNLS to compute projections onto the convex hull. However, DWNNLS and many other available sum-of-squares optimization solvers were not intended for usage with many variable problems, which result from the large training sets that are typical in machine learning applications. Thus, a new PROJECT subroutine is given, based on the highly customizable quadratic program solver BQPD. This solution is shown to be as robust as DELAUNAYSPARSE for projection onto both synthetic and real-world datasets, where other available solvers frequently fail. Although it is intended as an update for DELAUNAYSPARSE, due to the difficulty and prevalence of the problem, this solution is likely to be of external interest as well.

97 MATHEMATICS AND COMPUTING↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Learning nonlinear operators in latent spaces for real-time predictions of complex dynamics in physical systems

Abstract Predicting complex dynamics in physical applications governed by partial differential equations in real-time is nearly impossible with traditional numerical simulations due to high computational cost. Neural operators offer a solution by approximating mappings between infinite-dimensional Banach spaces, yet their performance degrades with system size and complexity. We propose an approach for learning neural operators in latent spaces, facilitating real-time predictions for highly nonlinear and multiscale systems on high-dimensional domains. Our method utilizes the deep operator network architecture on a low-dimensional latent space to efficiently approximate underlying operators. Demonstrations on material fracture, fluid flow prediction, and climate modeling highlight superior prediction accuracy and computational efficiency compared to existing methods. Notably, our approach enables approximating large-scale atmospheric flows with millions of degrees, enhancing weather and climate forecasts. Here we show that the proposed approach enables real-time predictions that can facilitate decision-making for a wide range of applications in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Direct quarkonium production in DIS from a joint CGC and NRQCD framework

We compute the differential cross section for direct quarkonium production in high-energy electron-nucleus collisions at small 𝑥. Our computation is performed within the nonrelativistic QCD factorization formalism that separates the calculation into short distance coefficients and long distance matrix elements that depend on the color and spin of the state. We obtain the short distance coefficients of the production of the heavy quark pair within the framework of the color glass condensate effective field theory, which resums coherent multiple interactions of the heavy quark pair with the nucleus to all orders. Our results are expressed as the convolution of perturbatively calculable functions with multipoint lightlike Wilson line correlators. In the correlation limit, we establish the correspondence between our color glass condensate formulation with calculations employing the transverse momentum dependent (TMD) framework. We extend this correspondence by resumming kinematic power corrections within the improved TMD framework, which interpolates between the TMD formalism and 𝑘 ⊥ -factorization formalism. We present a detailed numerical analysis, focusing on 𝐽/𝜓 production in the kinematics accessible at the future Electron-Ion Collider, highlighting the importance of genuine higher-order saturation contributions when the electron collides with a large nucleus. Our results are also valid in the photoproduction limit where we expect the largest contribution from genuine higher-order saturation contributions which could be accessed in ultraperipheral collisions of relativistic heavy ions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Aerodynamic Characterization of 3D Scanned Wind Turbine Blades Using Experimental and Computational Methods

This study presents an aerodynamic characterization of 3D scanned wind turbine blades using both experimental and computational methods. The research was conducted by Gulf Wind Technology and Sandia National Laboratories. The primary objective was to investigate the aerodynamic impacts of leading-edge manufacturing defects on wind turbine blades. The study utilized the Stratasys NEO 800 3D Printer for high-precision manufacturing and the GWT Accelerator Wind Tunnel for experimental testing. Computational simulations were performed using COMSOL Multiphysics to model the wind tunnel and analyze flow characteristics and OpenFOAM to study the aerodynamic impacts of leading-edge defects. OpenFAST was used to estimate how these defects can lead to revenue losses for wind farm operators as high as 6%. The results demonstrated significant aerodynamic performance variations due to defects, with detailed analysis provided through wind tunnel and CFD data. The findings contribute to the understanding of defect impacts on wind turbine blade performance and offer insights for future design improvements.

17 WIND ENERGY↗

Establishing a process-structure-property-performance framework for SLS additive manufacturing through integrated multiscale modeling

This study presents a comprehensive suite of high-fidelity computational models that integrate multiscale and multiphysics simulations to capture the full Selective Laser Sintering (SLS) additive manufacturing process—from initial melting and solidification to mechanical response under external loads. Process simulations are linked with mechanical analysis through Representative Volume Elements (RVEs), establishing a process-structure–property-performance framework. The interaction between laser light and polyamide 12 (PA12) powder is modeled, accounting for laser characteristics and the optical, thermal, and geometrical properties of the powder. The heat source is incorporated into a heat transfer model, coupled with crystallization kinetics and densification models to predict material density and crystallinity. The porosity distribution from the densification model and crystallinity interpolated from experimental data are used to construct the RVEs. A multi-mechanism constitutive model is then calibrated using mechanical tests to predict the stress–strain response. Simulation results show good agreement with experimental data in terms of porosity, crystallinity, and mechanical performance when sufficient laser power (62 W or higher) is used. This research supports the inverse design of 3D-printed structures by introducing a high-fidelity framework that combines multiscale and multiphysics modeling with experimental calibration for predictive and performance-driven additive manufacturing.

SLS↗

Equilipy: a python package for calculating phase equilibria

The CALPHAD (CALculation of PHAse Diagram) approach (Nigel Saunders & Miodownik, 1998) provides predictions for thermodynamically stable phases in multicomponent-multiphase materials across a wide range of temperatures. Consequently, the CALPHAD calculations became an essential tool in materials and process design (Luo, 2015). Such design tasks frequently require navigating a high-dimensional space due to multiple components involved in the system. This increasing complexity demands high-throughput CALPHAD calculations, especially in the rapidly evolving field of alloy design. In response to the need, we developed Equilipy an open-source Python package designed for calculating phase equilibria of multicomponent-multiphase systems. Equilipy is specifically tailored for high-throughput CALPHAD calculations, offering parallel computations across multiple processors and nodes with the given NPT input conditions namely elemental compositions (N), pressure (P), and temperature (T). Equilipy utilizes the program structure and Gibbs energy functions from the Fortran-based program, Thermochimica (Piro et al., 2013), with incorporating a new Gibbs energy minimization algorithm. This algorithm, originally developed by Capitani and Brown in 1987 (Capitani & Brown, 1987), has been revised and implemented to enhance the stability and performance of calculations. The Fortran codes are precompiled and interfaced with Python via F2PY, ensuring high computation speed. Benchmark tests shown in Figure 1 demonstrate that Equilipy’s computation speed is comparable to those of established commercial software, TC-Python and PanPython. This result highlights its efficiency and potential applications in various scientific and industrial fields.

97 MATHEMATICS AND COMPUTING↗

Operando microscopy for neuromorphic hardware

Microscopy techniques can uncover the physical properties and dynamic behaviours of materials, driving the discovery of emergent phenomena and guiding the design of next-generation computing hardware. As artificial intelligence becomes pervasive, the demand for high-performance materials to support sustainable information technologies is growing. Here, this Review highlights state-of-the-art imaging from electron and X-ray to optical techniques to probe the dynamics of neuromorphic materials, including operando characterization of devices. We examine design principles for neuromorphic materials, along with obstacles that hinder their development. Emphasis is placed on spatially and temporally resolved approaches that capture state changes including phase transitions, ferroic switching and spin-wave propagation that emulate biological components such as neurons, synapses and their connectivity. We discuss challenges in operando characterization and the integration of artificial intelligence-driven analysis for feedback-guided material discovery. Finally, we outline opportunities for real-time imaging of neuromorphic systems, paving the way towards adaptive, brain-inspired hardware.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Hybrid Cooling and Water Treatment for Resilient, Water-Self-Sufficient Data Centers

The continued growth of data center infrastructure is intensifying demand for freshwater resources, particularly in water-stressed regions, and is increasingly limiting sustainable capacity expansion. This work investigates a conceptual hybrid system that integrates freeze desalination with ultrasonic-assisted ice separation to enable on-site production of pure water from diverse sources, including seawater, brackish groundwater, and reclaimed industrial or agricultural wastewater. Simultaneously, the system produces low-temperature cooling streams that enhance heat removal in high power density computing environments. A process-level thermodynamic analysis is performed across a range of boundary and operating conditions, including variations in feed concentration, freezing temperature, and mass flow rate. The results are presented as performance curves relating feedwater concentration and target purified water output to the corresponding intake flow requirements, enabling estimation of source water demand per unit of Information Technology Equipment (ITE) energy consumption (kWh) across varying data center scales and operational scenarios. The corresponding electrical energy consumption for integrated cooling and water treatment is also evaluated as a function of system operating parameters and target water production levels. These results provide a basis for evaluating system feasibility across different conditions and for identifying parameter ranges in which integrated water treatment and cooling improve resource efficiency, thermal performance, and operational flexibility in data centers.

Elhefny, Aly [ORNL] (ORCID:0000000284907923)↗

Performance Testing of a Moving-Bed Gasifier Using Coal, Biomass, and Waste Plastic Blends with Washed and Unwashed Legacy Coals and Other Waste Fuels to Generate White Hydrogen

The objective of this effort, primarily funded by the United States Department of Energy (DOE), and led by the Electric Power Research Institute, Inc. (EPRI), with support by Hamilton Maurer International (HMI) and Sotacarbo S.p.A. (Sotacarbo), has been to qualify coal, biomass, and plastic waste blends based on performance testing of selected fuel pellet compositions in a pilot-scale updraft moving-bed (UDMB) gasifier. The testing provided relevant data to advance the commercial-scale design of the moving-bed gasifier to be able to successfully use these feedstocks to produce hydrogen. In particular, the effects of waste plastics on feedstock development (i.e., blending and pelletizing) and the resulting products (i.e., syngas compositions, organic condensate production, and ash characteristics) are the focus. The gasifier used for testing is HMI’s moving-bed gasifier, which has been proven capable of gasifying nearly all coal ranks. It has also shown the ability in prior testing work to gasify wood chips (biomass). However, mixtures of these fuels with plastic wastes have not been prepared and gasified together. The three feedstocks were densified and pelletized by California Pellet Mill (CPM) to meet the feedstock size required by Sotacarbo’s 30mm ID UDMB gasifier, under contract to HMI. The technical tasks and results from this two-year research project included: (1) Feed Procurement and Preparation: Nine different tri-fuel pellets were prepared from varying compositions of fresh mined PRB coal, corn stover biomass, and car fluff waste plastics. Tri-fuel pellets were produced by CPM and shipped to Sotacarbo’s test facility in Carbonia, Sardinia, Italy. (2) Test Plan Development: A test plan was created to define the test runs to be performed. The test plan detailed the different UDMB gasification tests to be performed in Sotacarbo’s 12-inch ID pilot scale gasifier, the process monitoring instrumentation used, and the extractive samples recovered for analysis of the total gasification process mass and energy balance. (3) Gasifier Testing: Nine different gasification runs were performed in the pilot-scale gasifier at Sotacarbo using nine different fuel feedstock compositions generated from varying mixtures of PRB coal, biomass, and plastic wastes. The testing generated performance data on gasification reaction efficiency and performance, yielding relevant data for models used to scale up the gasifier design. This task also included work to refurbish and reassemble the pilot gasifier at Sotacarbo and perform a baseline 100% PRB coal run. (4) Data Analysis and Reporting: Review of the data, determination of figures of merit, and interpretation of the results are reported in the project’s final report, published in March 2024. The results show that all tri-fuel pellets gasified well and maintained structural integrity throughout the gasification process. The syngas generated can be shifted to hydrogen by using commercial syngas shifting technologies. (5) High Fidelity computational fluid dynamics (CFD) Simulation: The National Energy Technology Laboratory (NETL) team performed CFD simulations of the UDMB gasifier for two of the tri-fuel pellets gasified in Sotacarbo’s pilot scale gasifier. The kinetic mechanisms for the pyrolysis of each constituent, PRB coal, corn stover biomass, and waste plastics are based on thermogravimetric analysis performed by Sotacarbo. The gasification model was validated by comparing the predicted syngas composition at the exit of the gasifier with the measured syngas composition. In addition, the reactor’s measured internal temperature profile agreed well with the predicted internal reactor temperature profile. These results validate that the model can be used to predict the performance of the updraft moving bed gasifier for different feedstocks and operating conditions. This paper summarizes the results of the completed work in which the pelletizing procedure was validated to ensure the viability of the tri-fuel pellets for the gasification runs performed at Sotacarbo’s 30 mm UDMB gasifier. The gasification performance data from this series of nine runs will enable modeling of a full-scale HMI industrial scale gasifier supporting both combined heat and power, and Hydrogen production from coal (both fresh mined and legacy) combined with various biomass and waste plastics. Additionally, plans and progress on a follow-up project, being executed by the same project team, will be presented. In this project, a total of twenty (20) different feedstocks are being prepared from varying compositions of biomass (both woody biomass and corn stover) with a mixture of legacy coal waste, plastic waste, and refuse-derived fuel (RDF). The testing will provide information on gasification reaction efficiency/performance, yielding relevant data for models used to scale up the gasifier design to 50 megawatt electric (MWe) (equivalent hydrogen production). Tests will also be performed on a bench-scale fluidized-bed gasifier for comparison purposes. The results of this testing will be used to specify the range of feedstock blends that can be successfully gasified as well as quantify gasifier outputs based on specific blends.

08 HYDROGEN↗

Porting ATLAS Fast Calorimeter Simulation to GPUs with Performance Portable Programming Models

FastCaloSim is a parameterized simulation of the particle energy response and of the energy distribution in the ATLAS calorimeter. It is a relatively small and self-contained package with massive inherent parallelism and captures the essence of GPU offloading via important operations like data transfer, memory initialization, floating point operations, and reduction. It was identified by the High Energy Physics Center for Computational Excellence project as a good testbed for evaluating the performance and ease of portability of programming models. In this paper, we will discuss the results of our evaluation of the porting process to Kokkos, SYCL, Alpaka, OpenMP and std::par (nvc++), and compare performance on NVIDIA, AMD and Intel GPUs, as well as multicore CPUs.

97 MATHEMATICS AND COMPUTING↗

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗