Search NASA⌕ Search

SEARCH · Search NASA

Results for “Performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.Multi

This work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Impact of Reordering on the LU Factorization Performance of Bordered Block-Diagonal Sparse Matrix

Power engineers rely on computer-based simulation tools to assess grid performance and ensure security. At the core of these tools are solvers for sparse linear equations. When transformed into a bordered block-diagonal (BBD) structure, part of the sparse linear equation solving can be parallelized. This work focuses on using the Schur-complement-based method for LU factorization on BBD matrices, specifically, Jacobian matrices from large-scale systems. Our findings show that the natural ordering method outperforms the default ordering method in computational performance for each block of the BBD matrix. This observation is validated using synthetic 25k-bus and 70k-bus cases, showing a speedup of up to 38% when using natural ordering without permutation. Additionally, the impact of the number of partitions is studied, and the result shows that computational performance improves with more, smaller partitions in the BBD matrices.

BBD matrix↗

Performance of a second generation of the Fermilab Constant Fraction Discriminator ASIC for time-stamping long-strip AC-LGAD sensors.

We present the design and performance characterization results of the second generation of the novel Fermilab Constant Fraction Discriminator ASIC (FCFD) developed to readout AC-coupled low gain avalanche detector (AC-LGAD) sensors. This study presents the performance of the ASIC which was optimized specifically for reading out strip AC-LGAD sensors designed for ePIC experiment at EiC. Performance was measured using charge injection and particle beams with prototype AC-LGAD sensors wirebonded to the FCFD ASIC.

Apresyan, A.↗

A Performance Model of In-Situ Techniques

The computational capacity of High-Performance Computing (HPC) systems increases continuously with the rapid development of central processing units (CPUs) and graphic processing units (GPUs), while the in-/output (IO) subsystem develops relatively slowly and storage capacity is also limited. Data-intensive applications, which are designed to leverage the high computational capacity of HPC resources, typically generate a considerable amount of data for post-processing visualizations and data analytics. The limited IO speed and storage space could lead to constraints in the actual performance of these applications and, therefore, scientific discovery. In-situ techniques, where data is visualized/analysed while still in memory rather than through disk, can contribute to alleviating these problems as they can reduce or even fully avoid data writing/reading through the IO subsystem to/from storage. However, the overall efficiency of insitu techniques crucially depends on the characteristics of both the in-situ tasks and the applications, and the resource distribution among them. Therefore, choosing the right in-situ approach (synchronous, asynchronous, or hybrid) and resource allocation is essential to minimize overhead and maximize the benefits of concurrent execution. In this paper, we present a performance model of in-situ techniques to find the most beneficial in-situ approach and the preferred resource configuration. We verify the high accuracy of our approach with over 6800 measurements and provide use cases with different applications.

Ju, Yi [Max Planck Computing and Data Facility, Ga↗

Thermal Performance of a Conduction-Cooled CCT Dipole ReBCO Magnet: Several Cycles of Cool-Down and Thermal Gradient Measurements

Here, this paper presents experimental results from conduction-cooled thermal testing of a ReBCO canted cosine theta (CCT) magnet (C2), originally designed and fabricated at LBNL using CORC cables. While the performance of the coil under liquid helium and nitrogen environments has been previously established, this study explores its behavior under conduction cooling using a large test cryostat at The Ohio State University. The magnet, measuring 613 mm in length and weighing ∼75 kg, consists of four helical layers wound with ReBCO-based CORC wire and was thermally anchored to a copper cold ring supported by a G-10 strongback. Cooling was provided by two Sumitomo RDK-415D cryocoolers, offering a combined 3 W at 4.2 K and 150 W at 77 K. Multiple thermal cycles were performed, with cooldown durations of up to 45 hours. Final base temperatures of approximately 10.8 K (at the coil edge) and 12.0 K (at the coil center) were achieved, with an axial temperature difference of approximately 1.2 K. The warm-up period extended over approximately 24.6 hours. Voltage measurements from all four layers were recorded during cooldown and warmup. The system demonstrated stable cooldown performance, repeatable gradients, and good thermal anchoring. These results support the feasibility of conduction cooling in large-scale HTS magnets, aligning with broader goals for “green” cryogen-free accelerator technologies and paving the way for more sustainable, scalable, and energy-efficient high-field magnet systems in next-generation particle accelerators.

Canted cosine theta magnets↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Thermal Performance of Triply Periodic Minimal Surface Lattice Structures in Single-Phase Dielectric Fluid Cooling of Power Electronics

Additive manufacturing has transformed thermal management by enabling the production of complex, optimized geometries that conventional manufacturing methods cannot achieve. This study investigates the single-phase convective heat transfer performance of gyroid triply periodic minimal surface (TPMS) lattice structures with functional porosity. TPMS structures provide high surface area to volume ratios and are amenable to 3D printing. A gyroid numerical model was created and validated against an existing experimental study with a similar feature size to the investigated geometries. The TPMS structure has a periodic width of 1.6 mm, a length of 10 mm, and a height of 4 mm, with a functional porosity ranging from 0.5 to 0.8, decreasing with distance from the heated surface. Three different flow configurations were examined for an inlet fluid temperature of 70 °C. The inlet velocities range from 0.01 to 1.2 m/s, corresponding to a Reynolds number range of 10–900 with a heat flux of 50 W/cm 2 applied at the base. AmpCool ® AC-110 dielectric fluid (Prandtl number 59.5) was used as the coolant. Thermal performance and friction characteristics were studied for the three flow orientations. The parallel flow configuration was identified as the most efficient for heat removal. A detailed analysis of the numerical results highlights the underlying physics behind the thermal performance differences among the flow configurations.

33 ADVANCED PROPULSION SYSTEMS↗

In-flight performance of the XRISM/Resolve detector system

The Resolve instrument was launched on board the XRISM observatory in early September 2023. The Resolve spectrometer is based on a high-sensitivity X-ray calorimeter detector system (DS) that has been successfully deployed in many ground and sub-orbital spectrometers. However, the Resolve instrument is the first long-term implementation in space. The instrument will provide essential diagnostics for nearly every class of X-ray emitting objects, from galactic supernova remnants to the outskirts of galaxy clusters, without degradation for spatially extended objects. The Resolve DS consists of a 36-pixel microcalorimeter array operated at a heat sink temperature of 50 mK. In pre-flight testing, the DS demonstrated a resolving power of better than 1300 at 6 keV with a simultaneous bandpass from below 0.3 keV to above 12 keV and a timing precision better than 100μs. An anti-coincidence detector placed directly behind the microcalorimeter array effectively suppresses background. The detector energy-resolution budget included terms for interference from the Resolve cooling system and the spacecraft. Additional terms for energy-scale stability, on-orbit effects, and use of mid-grade events were also included, predicting an end-of-life, on-orbit performance for high- and mid-resolution grade events that meet the requirement of 7 eV FWHM at 6 keV. Here, we discuss the actual on-orbit performance of the Resolve DS and compare this with the performance in pre-flight testing, on-orbit predictions, and the almost identical Hitomi/SXS instrument. We will also discuss the on-orbit gain stability, an assessment of on-orbit interference, and measurements of the on-orbit background.

Astronomy and AstroPhysics↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

Three Birds with One Stone: Improving Performance, Convergence, and System Throughput with NEST

Variational quantum algorithms (VQAs) have the potential to demonstrate quantum utility on near-term quantum computers. However, these algorithms often get executed on the highest-fidelity qubits and computers to achieve the best performance, causing low system throughput. Recent efforts have shown that VQAs can be run on low-fidelity qubits initially and high-fidelity qubits later on to still achieve good performance. We take this effort forward and show that carefully varying the qubit fidelity map of the VQA over its execution using our technique, Nest, does not just (1) improve performance (i.e., help achieve close to optimal results), but also (2) lead to faster convergence. We also use Nest to co-locate multiple VQAs concurrently on the same computer, thus (3) increasing the system throughput, and therefore, balancing and optimizing three conflicting metrics simultaneously.

qaoa↗

Correlation Between Microscopic Current Fluctuations Observed at Ultra-Microelectrodes and Macroscopic Bulk Electrolysis Performance in Redox-Active Microemulsions

Microemulsions (μEs) have been proposed as redox flow battery (RFB) electrolytes that maximize ionic conductivity and charge capacity by synergizing two immiscible phases. However, charge transfer during electrolysis in μEs is poorly understood. Here, we show that ultramicroelectrode electrolysis of ferrocene-loaded μEs –20%, 60%, and 90% water - reveals stochastic current fluctuations. These are differentiated in the scanning electrochemical microscopy (SECM) geometry, where power spectral density analysis showed distinct changes in the frequency contributions. SECM in the substrate generation-tip collection mode showed that fluctuations arise under mass-transfer control. Significant differences in the diffusion coefficient of ferrocene species were deducted from SECM approach curves, suggesting phase transfer behavior. Using bulk electrolysis, we calculated the charge accessibility and cycling behavior in the μEs. A decrease in the stochastic behavior of the μEs seems to correlate to a higher accessibility and cycling performance, with the 90% water μE displaying the best reversibility and the 60% the lowest. Altogether, these results suggest that Marangoni-type convection driven by concentration gradients and/or μE restructuring during charge transfer play a role in the electrochemical performance of μEs. This presents opportunities for screening and diagnosing the performance of these emerging RFB electrolytes.

25 ENERGY STORAGE↗

Impact of Module Configuration on Lithium-Ion Battery Performance and Degradation: Part I. Energy Throughput, Voltage Spread, and Current Distribution

Batteries are commonly connected in series and parallel to create modules that fulfill the power and energy requirements of specific applications. However, conclusions about battery performance and degradation under different conditions, as well as predictive models, are often derived from single cell cycling results. In this study, we evaluate the performance of six different series-parallel configurations of commercial lithium nickel manganese cobalt cells over hundreds of cycles. Each cell within the modules was individually instrumented for voltage, current, and temperature monitoring. We quantified the impact of module configuration on overall energy throughput, the voltage spread among series-connected cells, and the current heterogeneity in parallel-connected cells. This module cycling study, one of the broadest reported to date, supports systematic evaluation of the performance trade-offs, pack penalty, and safety implications of different module configurations.

25 ENERGY STORAGE↗

Effect of Carbon Support Heat Treatment on the Performance of Ion-Pair High-Temperature PEM Fuel Cells

Heat treatment can significantly alter the physical and chemical properties of carbon supports, thereby influencing the performance of proton exchange membrane (PEM) fuel cells. This study explores how varying carbon heat treatment temperatures—from 1,000 °C to 2,200 °C—affect the structure and performance of platinum-based catalysts in high-temperature PEM fuel cells. Increasing the heat treatment temperature led to a notable decrease in surface area, along with increases in carbon grain size and hydrophobicity. Although the catalyst supported on carbon treated at 1,000 °C exhibited the highest catalyst activity, the MEA using carbon treated at 1,500 °C delivered the best overall fuel cell performance. This is attributed to an optimized balance between hydrophobicity and accessible surface area, which enhances water management and catalyst utilization. These findings underscore the importance of carbon support engineering in improving the efficiency of ion-pair high-temperature PEM fuel cells.

08 HYDROGEN↗

Cobalt-Free Lithium and Manganese Rich Cathodes for Electric Vehicle Applications: Influence of Inherent Properties on Temperature-Dependent Performance

Lithium- and manganese-rich (LMR) oxides are attractive options as next-generation, earth-abundant cathodes. Despite inherent properties that result in less-than-ideal performance, advancements in this class of cathodes toward commercialization have been steady. Herein, we report a benchmark study on the influence these inherent properties exert on the discharge rate of Co-free, LMR//graphite (Gr) cells as a function of temperature. Because low-temperature performance is an essential consideration in the design of electric vehicle cells, understanding the limitations of current LMR systems is essential in understanding their current applicability as well as strategies for improvement. This work identifies two major limitations in the rate performance of LMR//Gr cells and suggests several areas of focus where practical strategies may provide improvements.

25 ENERGY STORAGE↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Battery Performance and Cost Model (BatPaC) Version 6.0

SF-26-016 The Battery Performance and Cost model (BatPaC) is a calculation method based on Microsoft Excel spreadsheets that has been developed at Argonne for estimating the performance and manufacturing cost of lithium-ion batteries for electric-drive vehicles including hybrid-electrics (HEV), plug-in hybrids (PHEVs) and pure electrics. BatPaC was first developed in 2007, was subsequently peer reviewed, and it has served Argonne researchers and the greater battery community in studying the impact of material properties on performance at the pack level. BatPaC has been updated and re-released multiple times since its original public release in 2011. This current version is BatPaC 6.0, which contains additional functionality needed to handle advances in automotive batteries, like the use of lithium metal and silicon anodes and the need to accommodate cell expansion and apply high levels of pressure.

KNEHR, KEVIN [Argonne National Laboratory (ANL), A↗

Lidar-Based Evaluation of HRRR Performance in California’s Diablo Range

The performance of the NOAA High-Resolution Rapid Refresh (HRRR) model for capturing low-level winds near a wind energy production site during summer 2019 is evaluated. This study catalogs the ability of HRRR to predict boundary layer dynamics relevant to wind energy interests over complex terrain, which has presented challenges for weather and energy forecasting. Performance is evaluated by comparing HRRR output to wind-profiling Doppler lidars at Lawrence Livermore National Laboratory Site 300. HRRR captured the diurnal profile of horizontal winds in the observed 150-m layer, despite strong underpredictions (∼4 m s −1 ) during evening and nighttime hours. These underpredictions may be a result of local speedup flows observed by the lidars, which were unresolved in HRRR due to their small spatial extent. HRRR bias magnitude relative to observations was found to be minimal during days with synoptic-scale troughs and strong 850-hPa geopotential gradients, while bias magnitude was maximal during days with synoptic ridging and weak 850-hPa geopotential gradients. To translate wind speed predictions to energy forecasting, generic turbine models were used to estimate power generation for turbines characteristic of the nearby Altamont Pass Wind Resource Area. Results show that HRRR-based energy estimates predicted daytime power generation adequately relative to lidar-based estimates with an 18-h lead time (bias magnitude < 0.4 MW from 0900 to 1400 LT) but overpredicted power during the rest of the diurnal cycle (bias > 1 MW). These results demonstrate conditions under which HRRR performs well for wind energy applications in complex terrain, while highlighting biases that require further investigation to support usage of a high-resolution model for wind energy forecasts.

Boundary layer↗