Search NASASearch

SEARCH · Search NASA

Results for “Computer systems performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING

Performance Evaluation of Low-Cost Condensers Immersed in the Water Tank of a Heat Pump Water Heater

This paper reports the effect of low-cost immersed condensers on the performance of a heat pump water heater (HPWH) that does not require a water-circulating pump. The proposed design is based on the immersion of an L-style condenser coil through an opening in the top of the water tank. The novel design aims to eliminate the need for the water-circulating pumps, thereby substantially improving efficiency and reducing costs and maintenance of the HPWH system. Comprehensive computational fluid dynamics (CFD) simulations and experimental tests were performed, and the results showed that the CFD simulations were close to the experimental data. Further results confirm that the L-style condenser can introduce buoyancy-driven flow and eliminate temperature stratification in the studied HPWH water tank, thereby improving the heat transfer coefficient and coefficient of performance of the HPWH. The testing results further showed that the L-style condenser enabled more than 53% higher coefficient of performance and 400% higher heat transfer coefficient compared with the straight vertical condenser, and achieved much faster water heating. Furthermore, the effects of immersed condensers with different circuits on the HPWH performance were evaluated and compared. All results indicate that the current energy-saving and inexpensive HPWH technology is technically valid, and the benefits of performance improvement and low cost are attractive in future HPWH design.

Gao, Zhiming [ORNL] (ORCID:0000000271397995)

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)

Ion-chain sympathetic cooling and gate dynamics

Sympathetic cooling is a technique often employed to mitigate motional heating in trapped-ion quantum computers. However, choosing system parameters such as number of coolants and cooling duty cycle for optimal gate performance requires evaluating trade-offs between motional errors and other slower errors such as qubit dephasing. The optimal parameters depend on cooling power, heating rate, and ion spacing in a particular system. In this study, we aim to analyze best practices for sympathetic cooling of long chains of trapped ions using analytical and computational methods. We use a case study to show that optimal cooling performance is achieved when coolants are placed at the center of the chain and provide a perturbative upper bound on the cooling limit of a mode given a particular set of cooling parameters. In addition, using computational tools, we analyze the trade-off between the number of coolant ions in a chain and the center-of-mass mode heating rate. We also show that cooling as often as possible when running a circuit is optimal when the qubit coherence time is otherwise long. These results provide a roadmap for how to choose sympathetic cooling parameters to maximize circuit performance in trapped-ion quantum computers using long chains of ions.

Cooling & trapping

PowderJet: Spherical metal powder production via multi-orifice droplet-on-demand metal jetting

Leading metal additive manufacturing techniques, such as laser powder bed fusion and directed energy deposition, rely on high-quality spherical metal powders. However, traditional powder production methods like gas atomization face limitations, including low in-spec yield, asphericity, and internal porosity. We introduce PowderJet, a powder production platform that uses electromagnetic pulses to eject liquid metal droplets from a multi-orifice nozzle. Unlike stochastic methods, PowderJet tightly controls powder size, distribution, and purity through a droplet-on-demand approach. We detail the system’s design, operation, and performance using a combined experimental and computational fluid dynamics (CFD) framework. Initial results with Al4008 and Cu110 alloys demonstrate successful production, yielding unsieved aluminum powder batches with a mean diameter of 200 µm and a narrow size distribution (15 µm standard deviation). The produced powders are highly spherical, achieving a roundness > 0.95. PowderJet operates with a small melt volume (3 mL) and supports continuous refilling, enabling production rates between 30 and 140 cm³/hr depending on jetting frequency, number of orifices and particle size. CFD simulations show that future systems could achieve rates exceeding 1000 cm³/hr for particle sizes as small as 40 µm. PowderJet’s high yield of in-spec powder makes it ideal for producing precious or hazardous materials that are inefficient to manufacture using conventional methods. This platform offers a scalable, precise, and efficient solution for producing high-quality powders tailored for advanced manufacturing applications.

Atomization

Atomistic‐Level Effects of Noncovalent Interactions and Crystalline Packing for Organic Material Structural Integrity upon Exposure to Gamma Radiation

Developing an atomistic understanding of ionizing radiation induced changes to organic materials is necessary for intentional design of greener and more sustainable materials for radiation shielding and detection. Cocrystals are promising for these purposes, but a detailed understanding of how the specific intermolecular interactions within the lattice upon exposure to radiation affect the structural stability of the organic crystalline material is unknown. This study evaluates atomistic-level effects of γ radiation on both single- and multicomponent organic crystalline materials and how specific noncovalent interactions and packing within the crystalline lattice enhance structural stability. Dose studies were performed on all crystalline systems and evaluated via experimental and computational methods. Changes in crystallinity were evaluated by p-XRD and free radical formation was analyzed via EPR spectroscopy. Type of intermolecular interactions and packing within the crystal lattice was delineated and related to the specific free radical species formed and the structural integrity of each material. Periodic DFT and HOMO-LUMO surface mapping calculations provided atomistic-level identifications of the most probable sites for the radicals formed upon exposure to γ radiation and relate intermolecular interactions and molecular packing within the crystalline lattice to experimental results.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Insights from Optimizing HPL Performance on Exascale Systems: A Comparative Analysis of Panel Factorization

High performance LINPACK (HPL) remains the primary benchmark for evaluating supercomputing performance. It includes many parts with substantial internal complexity, and its performance is affected by a large number of parameters that interact in ways that are difficult to predict on large-scale heterogeneous supercomputer systems. We present a comprehensive performance analysis of HPL on Frontier, the world’s first exascale supercomputer, which achieved HPL performance of 1.35 exaflops. Through empirical parameter tuning, detailed modeling, and comparative evaluation, we uncover critical performance insights, share lessons learned, and outline best practices for effective parameter tuning on exascale systems. We introduce and evaluate two novel PDFACT strategies: a dedicated-thread (DT) variant and a GPU-based variant (GPUPDFACT) implementation using HIP cooperative groups, demonstrating that GPU-based factorization outperforms conventional CPU-based PDFACT on Frontier’s architecture. Our findings establish key performance factors for HPL on exascale systems and offer valuable guidance for future high-performance computing and benchmarking efforts.

Lu, Hao [ORNL] (ORCID:000000018941870X)

Advanced CO 2 Capture Solvent Systems for Dynamic Power Generation: Quarterly Research Performance Progress Report, QR4 (Q4FY24)

We developed an integrated Computational Fluid Dynamics (CFD) model to simulate the multi-physics coupled cooling process of mixed gas by cold water within a Direct Contact Cooler (DCC) equipped with a rotating packing bed (RPB). The model captures the interactions between fluid dynamics, heat transfer, mass transport, and phase transitions, while accounting for key operational variables such as RPB rotational speed and the mass flow rates of both liquid and gas. The CFD model has been validated using experimental data, specifically by comparing predicted outflow gas and liquid temperatures to measured results. Our findings demonstrate the significant effects of RPB rotational speed and mass flow rates on cooling performance, providing valuable insights for optimizing DCC efficiency in industrial applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Linear Solver for Electromagnetic Simulation of General Distribution Feeders

High-fidelity electromagnetic transient (EMT) modeling is required for accurate simulation and analysis of power system dynamics in modern distribution feeders. However, the high-fidelity of EMT models often leads to significant computational challenges, particularly in terms of computational resources and simulation time. This paper investigates the development and application of a detailed EMT model for general distribution feeders, with a focus on improving computational efficiency. A direct linear solver is proposed for a bordered block diagonal (BBD) matrix structure commonly encountered in a EMT model of distribution feeders. The solver integrates the Schur complement method with the block tridiagonal matrix algorithm to enhance the computational performance. The proposed solver is validated using the primary feeder of the IEEE 342-node test system, demonstrating its accuracy and efficiency in EMT simulations. Furthermore, the solver’s performance is benchmarked against MATLAB’s built-in linear solvers, showing significant improvements in computation time while maintaining high fidelity and accuracy in simulation results.

Choi, Jongchan [ORNL] (ORCID:000000025952455X)

Enhanced climate reproducibility testing with false discovery rate correction

Simulating the Earth's climate is an important and complex problem, thus climate models are similarly complex, comprised of millions of lines of code. In order to appropriately utilize the latest computational and software infrastructure advancements in Earth system models running on modern hybrid computing architectures to improve their performance, precision, accuracy, or all three; it is important to ensure that model simulations are repeatable and robust. This introduces the need for establishing statistical or non-bit-for-bit reproducibility, since bit-for-bit reproducibility may not always be achievable. Here, we propose a short-simulation ensemble-based test for an atmosphere model to evaluate the null hypothesis that modified model results are statistically equivalent to that of the original model. We implement this test in version 2 of the US Department of Energy's Energy Exascale Earth System Model (E3SM). The test evaluates a standard set of output variables across the two simulation ensembles and uses a false discovery rate correction to account for multiple testing. The false positive rates of the test are examined using re-sampling techniques on large simulation ensembles and are found to be lower than the currently implemented bootstrapping-based testing approach in E3SM. We also evaluate the statistical power of the test using perturbed simulation ensemble suites, each with a progressively larger magnitude of change to a tuning parameter. The new test is generally found to exhibit more statistical power than the current approach, being able to detect smaller changes in parameter values with higher confidence.

Kelleher, Michael E. [Oak Ridge National Laborator

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL

Elastic Stochastic Full Waveform Inversion (eSFWI)

This collaboration between Lawrence Livermore National Security, LLC (LLNS) as manager and operator of Lawrence Livermore National Laboratory (LLNL) and Chevron USA Inc., acting through its Chevron Technical Center division, aimed at developing next-generation computational methods for the Elastic Stochastic Full Waveform Inversion (eSFWI). Seismic imaging is heavily used in the oil and gas industry for identifying and operating subsurface reservoirs. Improved seismic imaging methods can improve productivity, lower costs, and improve operational and environmental safety. This CRADA demonstrated that new high-performance computing (HPC) architectures being rolled out over the next five years can enable unprecedented seismic imaging resolution when using eSFWI techniques to process active seismic data. An open-source computational mini-application was developed, capable of demonstrating near-peak performance for eSFWI algorithms on CPU and GPU enabled HPC platforms. Performance was demonstrated on LLNL HPC systems such as Lassen, as well as on Chevron systems. This project benefited Chevron USA Inc. by demonstrating the potential computational efficiency of their full waveform inversion capabilities used to characterize oil/gas reservoirs, which in turn benefits the public through potential increases in capabilities to perform analysis of leasing sites.

04 OIL SHALES AND TAR SANDS

Simplified projection on total spin zero for state preparation on quantum computers

Here, we introduce a simple algorithm for projecting on J = 0 states of a many-body system by performing a series of rotations to remove states with angular momentum projections greater than zero. Existing methods rely on unitary evolution with the two-body operator J 2 , which when expressed in the computational basis contains many complicated Pauli strings requiring Trotterization and leading to very deep quantum circuits. Our approach performs the necessary projections using the one-body operators J x and J z . By leveraging the method of Cartan decomposition, the unitary transformations that perform the projection can be parametrized as a product of a small number of two-qubit rotations, with angles determined by an efficient classical optimization. Given the reduced complexity in terms of gates, this approach can be used to prepare approximate ground states of even-even nuclei by projecting onto the J = 0 component of deformed Hartree-Fock states. We estimate the resource requirements in terms of the universal gate set {H,S, CNOT ,T} and briefly discuss a variant of the algorithm that projects onto J = 1/2 states of a system with an odd number of fermions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Machine learning without a processor: Emergent learning in a nonlinear analog network

Standard deep learning algorithms require differentiating large nonlinear networks, a process that is slow and power-hungry. Electronic contrastive local learning networks (CLLNs) offer potentially fast, efficient, and fault-tolerant hardware for analog machine learning, but existing implementations are linear, severely limiting their capabilities. These systems differ significantly from artificial neural networks as well as the brain, so the feasibility and utility of incorporating nonlinear elements have not been explored. Here, we introduce a nonlinear CLLN—an analog electronic network made of self-adjusting nonlinear resistive elements based on transistors. We demonstrate that the system learns tasks unachievable in linear systems, including XOR (exclusive or) and nonlinear regression, without a computer. We find our decentralized system reduces modes of training error in order (mean, slope, curvature), similar to spectral bias in artificial neural networks. The circuitry is robust to damage, retrainable in seconds, and performs learned tasks in microseconds while dissipating only picojoules of energy across each transistor. This suggests enormous potential for fast, low-power computing in edge systems like sensors, robotic controllers, and medical devices, as well as manufacturability at scale for performing and studying emergent learning.

Science & Technology - Other Topics

Enhancing ACPF Analysis: Integrating Newton-Raphson Method with Gradient Descent and Computational Graphs

This paper presents a new method for enhancing Alternating Current Power Flow (ACPF) analysis. The method integrates the Newton-Raphson (NR) method with Enhanced-Gradient Descent (GD) and computational graphs. The integration of renewable energy sources in power systems introduces variability and unpredictability, and this method addresses these challenges. It leverages the robustness of NR for accurate approximations and the flexibility of GD for handling variable conditions, all without requiring Jacobian matrix inversion. Furthermore, computational graphs provide a structured and visual framework that simplifies and systematizes the application of these methods. The goal of this fusion is to overcome the limitations of traditional ACPF methods and improve the resilience, adaptability, and efficiency of modern power grid analyses. We validate the effectiveness of our advanced algorithm through comprehensive testing on established IEEE benchmark systems. Furthermore, our findings demonstrate that our approach not only speeds up the convergence process but also ensures consistent performance across diverse system states, representing a significant advancement in power flow computation.

24 POWER TRANSMISSION AND DISTRIBUTION