Search NASA⌕ Search

SEARCH · Search NASA

Results for “Structure optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

An Optimized Parameterization of Sub‐Grid Scale Advection for Convection Permitting Models

Convection‐permitting models (CPMs) explicitly resolve deep convection yet under‐resolve the organized lateral exchanges among drafts and their environment that control entrainment/detrainment, precipitation efficiency, and mesoscale structure. In this work, we introduce the Optimized Advection Scheme (OAS), which introduces a small rotation of the Cartesian frame of reference for the horizontal winds relative to other variables used in advection that induces cross‐gradient transport to mimic under‐resolved convective mixing. The rotation angle is selected to minimize the Kullback–Leibler divergence between the simulated and satellite observed precipitation intensity distributions, yielding a physically consistent perturbation that is computationally inexpensive and portable. Optimized Advection Scheme is implemented in WRF and evaluated over Amazon (April 2014). It shifts precipitation–precipitable‐water joint distributions toward lighter rain, reduces overly intense rates, and improves mesoscale convective system (MCS) lifetime and propagation. Mechanistically, the added cross‐gradient transport promotes convective detrainment and environmental mixing, which cools and moistens the mid‐troposphere, weakens downward momentum transport, alleviates excessive downwelling shortwave biases, and warms the surface temperature. The optimized rotation angle yields comparable improvements at 4‐km and 1‐km grid spacing, demonstrating resolution‐independent benefits across the CPM gray zone. By targeting the dynamical root of under‐mixed convective circulations, rather than tuning model microphysics or closures, OAS delivers robust, scale‐aware improvements in precipitation statistics, cloud vertical structure, and characteristics of MCS (MCSs), offering a practical pathway to more reliable CPM simulations for weather and climate applications.

CPM↗

Atomic structure of different surface terminations of polycrystalline ZnPd

The intermetallic compound ZnPd has been found to have desirable characteristics as a catalyst for the steam reforming of methanol. The understanding of the surface structure of ZnPd is important to optimize its catalytic behavior. However, due to the lack of bulk single-crystal samples and the complexity of characterizing surface properties in the available polycrystalline samples using common experimental techniques, all previous surface science studies of this compound have been performed on surface alloy samples formed through thin-film deposition. In this study, we present findings on the chemical and atomic structure of the surfaces of bulk polycrystalline ZnPd studied by a variety of complementary experimental techniques, including scanning tunneling microscopy (STM), x-ray photoelectron spectroscopy (XPS), low energy electron microscopy (LEEM), photoemission electron microscopy (PEEM), and microspot low-energy electron diffraction ( μ -LEED). These experimental techniques, combined with density functional theory (DFT)-based thermodynamic calculations of surface free energy and detachment kinetics at the step edges, confirm that surfaces terminated by atomic layers composed of both Zn and Pd atoms are more stable than those terminated by only Zn or Pd layers. DFT calculations also demonstrate that the primary contribution to the tunneling current arises from Pd atoms, in agreement with the STM results. The formation of intermetallics at surfaces may contribute to the superior catalyst properties of ZnPd over Zn or Pd elemental counterparts. Published by the American Physical Society 2024

36 MATERIALS SCIENCE↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

Ionic-content-driven restructuring of spirobisindane ionene networks: implications for mechanics, self-healing, and gas transport

Polymers of intrinsic microporosity (PIMs) offer exceptional gas permeability but remain brittle and susceptible to physical aging, limiting their durability in separation applications. Here, we introduce a reconfigurable microporous polymer network that uniquely integrates permanent PIM microporosity with autonomous, intrinsic self-healing driven by imidazolium-based ionic motifs. Spirobisindane units generate the intrinsic free-volume architecture, while an imidazolium-containing polyamide ionene supplies dynamic ionic and hydrogen-bonding interactions that reorganize under mild activation. Incorporation of imidazolium-based ionic liquids further tunes cohesion, mobility, and densification, enabling the network to relax, re-associate, and retain microporosity without structural collapse. Through a comprehensive multiscale approach combining spectroscopy, scattering, thermal and mechanical characterization with all-atom molecular dynamics and density functional theory calculations, we elucidate how ionic content, as a single control parameter that reshapes free-volume distributions, modulates local coordination environments, and governs relaxation and healing kinetics. At intermediate ionic loadings, the networks achieve rapid, repeatable self-healing while maintaining CO$_2$ selectivity, demonstrating an optimal balance between segmental mobility and structural integrity. By establishing how hierarchical ionic interactions couple structure, dynamics, and transport in microporous ionene networks, this work provides generalizable design rules for adaptive soft-matter systems that require simultaneous mechanical resilience, reconfigurability, and selective gas transport.

36 MATERIALS SCIENCE↗

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S↗

Conceptual Structural Design and Analysis of a 20 T Hybrid Cos Dipole for Future Particle Colliders

Here, to reach high collision energy for future high-energy particle colliders, like the Future Circular Collider (FCC) or the Muon Collider, it is required to achieve high field strength of the bending dipoles. Currently, the practical limit for Nb$_{\text{3}}$ Sn technology is around 16 T and, in order to further increase the magnetic field, the superconducting magnet community is considering High Temperature Superconductors (HTS), in particular Bi-2212 and REBCO conductors. However, their relevant higher cost has led the community to consider a hybrid approach where HTS materials are used in the high field region of the coils with so-called insert coils, and Low Temperature Superconductors are involved in the lower field part ($< $ 16 T) with so-called outsert coils. This paper describes the conceptual mechanical design of a 20 T hybrid cos$\theta$ dipole configuration. The high stress levels that the structure is facing due to the high magnetic field are discussed. Moreover, it presents the results of the optimization analysis of the shell-based support structure based on the key-and-bladder technology that provides the azimuthal pre-stress during room temperature assembly and cooldown to cryogenic temperatures. The aim of this work is to present a feasible design that satisfies the stress requirements.

D'Addazio, Marika [Politecnico di Torino (Italy); ↗

ECP libraries and tools: An overview

The Exascale Computing Project (ECP) Software Technology and Co-Design teams addressed the growing complexities in high-performance computing (HPC) by developing scalable software libraries and tools that leverage exascale system capabilities. As we enter the exascale era, the need for reusable, optimized software solutions that can handle the unique challenges posed by these systems becomes increasingly important. The primary challenges the ECP teams faced were to create software libraries and tools that are performant on exascale architectures and portable and usable across diverse hardware platforms. Efforts addressed issues related to concurrent execution, memory management, and the integration of heterogeneous computing resources, such as GPUs from multiple vendors. The ECP’s strategy involved a structured development process encompassing the creation, optimization, and deployment of software in collaboration with industry, academia, and national laboratories. The project was organized into several technical areas: co-design of domain-specific suites with target applications, programming models and runtimes, development tools, mathematical libraries, data and visualization tools, and software ecosystem and delivery mechanisms. ECP has successfully developed a large portfolio of software libraries and tools that demonstrate significant improvements in performance and scalability on exascale systems. These products have been integrated into the Department of Energy’s computing facilities, supporting various scientific applications and ensuring robust performance across different hardware setups. ECP advancements in software development for exascale computing highlight the importance of a collaborative and adaptive approach to handling next-generation HPC systems complexities. The lessons learned emphasize the need for continuous engagement with end-users and vendors, and the importance of maintaining a balance between innovation and practical implementation. Future efforts will focus on ensuring scalability, keeping pace with rapid hardware advancements, and further enhancing the interoperability and usability of the software ecosystem. In conclusion, subsequent articles in this special issue provide in-depth discussions and case studies into specific library and tool efforts.

97 MATHEMATICS AND COMPUTING↗

Direct Air Capture Using Trapped Small Amines in Hierarchical Nanoporous Capsules on Porous Electrospun Fibers (Final Technical Report)

This report summarizes the carbon capture research and development conducted by The State University of New York at Buffalo (UB) and GTI Energy (GTI) for award “DE-FE0031969: Direct Air Capture Using Trapped Small Amines in Hierarchical Nanoporous Capsules on Porous Electrospun Fibers” sponsored by the U.S. Department of Energy (DOE). The objective of this project is to develop an innovative sorbent structure of trapped small amines in HNC embedded in PEF for DAC. This involves tailoring both sorbent and PEF materials to achieve a compact system for DAC with high capacity for CO 2 at concentrations typically available in air and at near ambient conditions. An innovative sorbent structure of trapped small amines in hierarchical nanoporous capsules (HNC) embedded in porous electrospun fibers (PEF) was developed for direct air capture (DAC). This involves tailoring both sorbent and PEF materials to achieve a compact system for DAC with high capacity for CO 2 at concentrations typically available in air and at near ambient conditions. An interfacial polymerization process was developed, which utilized loaded amines inside mesoporous silica and trimesoyl chloride (TMC) dissolved in organic solvents as the precursors, to generate a polyamide (PA) coating layer on mesoporous silica and thus trap amines. Reaction conditions, including TMC concentration, organic solvents, reaction time, etc., for interfacial polymerization were optimized to effectively trap loaded amines, and cyclic heating-cooling operation was conducted to evaluate the coating quality. Larger pore volume mesoporous silica was also synthesized to increase amine loading and thus increase CO 2 capacity. The optimized sorbent material exhibited CO 2 capacity as high as 4.88 mmol/g under humid DAC conditions and negligible loss (<1%) during 10 cyclic heating-cooling operations. The optimized PA-coated sorbent also showed fast adsorption and desorption kinetics, with <20% t1/2 increase compared to uncoated sorbent. PEF fabrication conditions, including organic solvents for dissolving core and shell polymers, voltage, distance from the nozzle to the collection panel, etc. were adjusted to better incorporate HNC. After incorporating the optimized sorbent material into PEF, the structured sorbent had a CO 2 capacity of approximately 4.0 mmol/g under humid DAC conditions, with capacity loss of 0.17% per cycle and t1/2 increase less than 10%. A techno-economic analysis (TEA) for the process design for a DAC system based on our developed sorbent structure of trapped small amines in HNC embedded in PEF was conducted. The process design included process description and major equipment sizing and energy and mass balances in addition to scale-up research results and estimated capture cost. Aspen Adsorption Simulator was used to fit the experimentally measured breakthrough curves and extract equilibrium and kinetic data of the optimized sorbent. Our results indicated that for a DAC plant with CO 2 productivity of 3,000 tonne/year, the levelized cost of CO 2 capture was $\$$612/tonne, with the largest contribution of 44.33% from the fixed operation cost. Increasing CO 2 productivity, while maintaining similar fixed operation cost, is expected to significantly reduce the CO 2 capture cost. A sensitivity study was also conducted to understand the influence of total plant cost, sorbent cost, CO 2 concentration in the feed, sorbent mat lifetime, sorbent regeneration electricity, and adsorption blower pressure drop on the levelized cost of CO 2 capture, revealing a capture cost range of $\$$520-870/tonne.

36 MATERIALS SCIENCE↗

Efficient Parameterization of Density Functional Tight-Binding for 5 f -Elements: A Th–O Case Study

Density functional tight binding (DFTB) models for f-element species are challenging to parametrize owing to the large number of adjustable parameters. The explicit optimization of the terms entering the semiempirical DFTB Hamiltonian related to f orbitals is crucial to generating a reliable parametrization for f-block elements, because they play import roles in bonding interactions. However, since the number of parameters grows quadratically with the number of orbitals, the computational cost for parameter optimization is much more expensive for the f-elements than for the main group elements. In this work we present a set of efficient approaches for mitigating the hurdle imposed by the large size of the parameter space. A novel group-by-orbital correction functions for two-center bond integrals was developed. With this approach the number of parameters is reduced, and it grows linearly with the number of elements, maintaining the accuracy and the number of parameters, in the case of f elements, by more than 40%. The parameter optimization step was accelerated by means of the mini-batch BFGS method. This method allows parameter optimizations with much larger training sets than other single batch methods. A stochastic optimizer was employed that helped overcome shallow local minima in the objective function. The proposed algorithm was used to parametrize the DFTB Hamiltonian for the Th–O system, which was subsequently applied to the study of ThO 2 nanoparticles. The training set consisted of 6322 unique structures, which is barely feasible with conventional optimization methods. The optimized parameter set, LANL-ThO, displays good agreement with DFT-calculated properties such as energies, forces, and structures for both clusters and bulk ThO 2 . Benefiting from the fewer number of parameters and lower computational costs for objective function evaluations, this new approach shows its potential applications in DFTB parametrization for elements with high angular momentum, which present a challenge to conventional methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TMD phenomenology motivated by nonperturbative structures

This talk summarized work done recently to organize the steps for implementing TMD phenomenology in a way optimized for contexts where the extraction and interpretation of hadronic structures and nonperturbative effects is the primary driving motivation.

Rogers, Ted↗

Elucidating the Interfacial Effects of Nonmetallic Elements on the Dehydrogenation Behavior of Nanoconfined NaAlH 4 in Zeolite-Templated Carbon

Confining materials within nanoscale volumes alters their physical and chemical properties, with positive consequences for energy storage, conversion, and catalysis. The pore structure and composition of scaffolds are essential variables for optimizing these properties, with carbon-based materials being preferred due to their tunable porous structures and chemical versatility. This study investigates the influence of surface functional groups on the dehydrogenation kinetics of nanoconfined NaAlH4 using zeolite-templated carbons (ZTCs). Here we focus on oxygen functional groups commonly present as intrinsic impurities on carbon scaffolds, analyzing three ZTC scaffolds to determine how their concentrations and configurations affect dehydrogenation behavior. Our findings reveal that carbonyl groups enhance charge transfer and destabilize Al–H bonds more effectively than ether or phenol groups. This indicates that the type of oxygen functional group is more critical than the quantity, highlighting the importance of properly tailoring oxygen defects to improve hydrogen storage performance in nanoconfined systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ternary Phosphides Ba M 2 P 2 : Tailoring Crystal and Electronic Structures Enables Highly Efficient HER Electrocatalysis

Binary transition metal phosphides and their solid solutions have emerged as promising hydrogen evolution reaction (HER) catalysts. Although many research endeavors have adopted strategies to vary compositions to optimize catalytic performance, they mainly focus on binary structures, which represent only a small fraction of the abundant phase space of structure types among transition metal phosphides. Here, the largely unexplored class of ternary and multinary ordered phosphides in catalysis comprises two or more metals with quite different chemical nature, concealing the structure–property relationships essential for advancing catalyst design. Here, we explored phosphides crystallizing in one of the most abundant ordered intermetallic structure types, —the ThCr 2 Si 2 type, —where square nets of 3d transition metal M and P atoms are separated by layers of electropositive Ba cations. Four ternary BaM 2 P 2 (M = Fe, Fe/Cu, Fe/Ni, Ni) catalysts were synthesized and characterized. BaNi 2 P 2 showed high HER activity in acidic electrolyte, which required an overpotential, η 10 , of only 62 mV to drive current density j = –10 mA/cm 2 and high stability with a potential drop rate of 0.25 mV/h. BaNi 2 P 2 outperformed other Ni-based catalysts, such as Ni 2 P and Ni 5 P 4 . Notably, at current densities above –170 mA/cm 2 , BaNi 2 P 2 outperformed the standard Pt electrode measured under identical conditions. Electronic structure analysis revealed a volcano-type activity trend among the four BaM 2 P 2 catalysts based on their d-band center positions, highlighting the role of electropositive Ba cations in shifting the Ni-3d orbitals into an optimal position.

BaNi2P2↗

Designing Cellular Metal Structures for Thermal Insulation

This project focused on developing topology optimization software to design advanced metal thermal insulators. Initially, solid designs were created that matched the thermal performance of current baseline designs but were significantly heavier. To address this, cellular materials were incorporated, specifically the octet structure which is known for its high strength-to-weight ratio and thermal properties. By leveraging these cellular designs at various densities, superior thermal and mechanical performance was achieved without added weight. This novel approach enhances thermal management and structural integrity under extreme conditions, offering promising advancements for thermal protection systems.

36 MATERIALS SCIENCE↗

Uncovering Structure–Conductivity Relationships in Anion Exchange Membranes (AEMs) Using Interpretable Machine Learning

Anion exchange membranes (AEMs) play a vital role in the performance of water electrolyzers and fuel cells, yet their discovery and optimization remain challenging due to the complexity of structure–property relationships. In this study, we introduce a machine learning framework that leverages conditional graph neural networks (cGNNs) and descriptor-based models and a hybrid graph neural network (HGARE) to predict and interpret ionic conductivity. The descriptor-based pipeline employs principal component analysis (PCA), ablation, and SHAP analysis to identify factors governing anion conductivity, revealing electronic, topological, and compositional descriptors as key contributors. Beyond prediction, dimensionality reduction and clustering are performed by employing t-SNE and KMeans as well as SOM, which reveal distinct membranes clusters, some of which were enriched with high anion conductivity. Among graph-based approaches, the graph convolutional (GCN) achieved strong predictive performance, while the Hybrid Graph Autoencoder-Regressor Ensemble (HGARE) achieved the highest accuracy. Additionally, atom-level saliency maps from GCN provide spatial explanations for conductive behavior, revealing the importance of polarizable and flexible regions. This work contributes to the accelerated and data-driven design of high-performance AEMs.

Naghshnejad, Pegah [Department of Chemical Enginee↗

Manipulating Na/TM Ratio‐Driven Structural Heterogeneity of O3‐NaNi 1/3 Fe 1/3 Mn 1/3 O 2 Cathode for High‐Voltage Sodium‐Ion Batteries

The stability of O3-type NaNi 1/3 Fe 1/3 Mn 1/3 O 2 under high-voltage cycling is dictated by how synthesis encodes lattice strain and redox heterogeneity. Here, in this study, the role of Na:TM stoichiometry is systematically resolved by tuning the NaOH:precursor ratio during solid-state synthesis. The stoichiometric condition (Na:TM = 1.00) yields minimized microstrain, enabling uniform O3–P3 phase evolution and homogeneous multi-metal redox with preserved octahedral symmetry. In contrast, Na-excess compositions inherit disordered intermediates and heterogeneous distortion fields that trigger abrupt multiphase transitions and promote localized charge redistribution. In situ XRD captures the divergence in phase-transition pathways, TXM resolves particle-level redox heterogeneity, and XANES corroborates a stronger and more reversible Fe redox contribution at stoichiometry, shifting to diminished Fe participation and spatially inhomogeneous redox at higher Na content. These results establish Na:TM stoichiometry as a critical synthesis parameter controlling both structural coherence and redox stability. Electrochemically, the stoichiometric composition exhibits smooth voltage profiles with minimal polarization growth and retains nearly 80% of its initial capacity after 100 cycles even at an extended 4.2 V cutoff, whereas Na-excess compositions show significantly reduced initial coulombic efficiency and rapid voltage fade. Precise stoichiometric tuning provides a scalable route to defect-suppressed O3 frameworks, enabling structurally resilient, high-voltage sodium-layered cathodes.

36 MATERIALS SCIENCE↗

AIMD‐Based Protocols for Modeling Exciplex Fluorescence Spectra and Inter‐System Crossing in Photocatalytic Chromophores

ABSTRACT This study introduces a computational protocol for modeling the emission spectra of exciplexes using excited‐state ab initio molecular dynamics (AIMD) simulations. The protocol is applied to a model exciplex formed by oligo‐p‐phenylenes (OPPs) and triethylamine (TEA), which is of interest in the context of photocatalytic reduction of . AIMD facilitates efficient sampling of the conformational space of OPP3 and OPP4 exciplexes with TEA, offering a dynamic alternative to previously employed static methods. The AIMD‐based protocol successfully reproduces experimental emission spectra for OPP‐TEA exciplexes, agreeing with previous computational and experimental findings. The results show that AIMD simulations provide an efficient means of sampling the conformational space of these exciplexes, requiring less user input and, in some instances, fewer computational resources than multiple excited‐state optimizations initiated from user‐specified initial structures. The study also evaluates the yield of intersystem crossing (ISC) using AIMD and Landau‐Zener probability. The results suggest that ISC is a minor decay channel for OPP3 and OPP4. This work provides new insights into the structural flexibility and emission characteristics of OPP‐TEA photoredox catalyst systems, potentially contributing to improved design strategies for organic chromophores in reduction applications.

Giudetti, Goran [Department of Chemistry Universit↗

Kinetic control of phase evolution and defect-mediated recombination in AACVD-grown copper antimony sulphide thin films

Copper antimony sulphide (CAS) is a multinary chalcogenide semiconductor in which small deviations from stoichiometry can drive phase competition and strong defect-mediated modulation of optoelectronic properties. However, the systematic roles of copper precursor fraction and deposition time in governing phase evolution, off-stoichiometry, and recombination dynamics in aerosol-assisted chemical vapour deposition (AACVD)-grown undoped CAS thin films remain insufficiently understood. In this work, undoped CAS thin films were deposited by AACVD using Cu(dedtc)2 and Sb(dedtc)3 single-source precursors at 550 °C and a carrier gas flow rate of 150 sccm, while the Cu(dedtc)2 mole fraction (x = 0.15 – 0.55) and deposition time (1 – 2 hours) were systematically varied to probe how growth kinetics influence phase composition, microstructure, and defect-mediated optical properties without post-deposition annealing or extrinsic doping. Increasing copper precursor content drives phase evolution toward tetrahedrite-dominant CAS films at intermediate Cu(dedtc)2 mole fractions, with the film deposited at x = 0.35 exhibiting the strongest tetrahedrite character within the parameter space examined. The films are also copper-rich, antimony-poor, and sulphur-deficient, consistent with off-stoichiometric growth and intrinsic defect formation, plausibly including copper interstitials, Cu-on-Sb antisites, and sulphur vacancies. These growth-dependent compositional deviations are accompanied by tunable indirect optical bandgaps of approximately 1.60 – 2.18 eV and weak visible photoluminescence governed by defect-mediated recombination. Time-resolved photoluminescence reveals bi-exponential decay behaviour with lifetimes of approximately 0.1 – 3.6 ns, with emission dominated by slower donor–acceptor pair recombination and a smaller contribution from faster trap-assisted pathways. Collectively, these results establish an explicit kinetic process–structure–defect–property relationship for AACVD-grown CAS thin films and provide a growth–structure–property framework relevant to future optimization of CAS-based optoelectronic and energy materials.

36 MATERIALS SCIENCE↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗