Search NASA⌕ Search

SEARCH · Search NASA

Results for “high performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Effects of renormalon scheme and perturbative scale choices on determinations of the strong coupling from e + e − event shapes

We study the role of renormalon cancellation schemes and perturbative scale choices in extractions of the strong coupling constant α s ( m Z ) and the leading nonperturbative shift parameter Ω 1 from resummed predictions of the e + e − event shape thrust. We calculate the thrust distribution to N L 3 L ′ resummed accuracy in soft-collinear effective theory (SCET) matched to the fixed-order O ( α s 2 ) prediction, and perform a new high-statistics computation of the O ( α s 3 ) matching in , although we do not include the latter in our final α s fits due to some observed systematics that require further investigation. We are primarily interested in testing the phenomenological impact sourced from varying amongst three renormalon cancellation schemes and two sets of perturbative scale profile choices. We then perform a global fit to available data spanning center-of-mass energies between 35–207 GeV in each scenario. Relevant subsets of our results are consistent with prior SCET-based extractions of α s ( m Z ) , but we are also led to a number of novel observations. Notably, we find that the combined effect of altering the renormalon cancellation scheme and profile parameters can lead to few-percent-level impacts on the extracted values in the α s − Ω 1 plane, indicating a potentially important systematic theory uncertainty that should be accounted for. We also observe that fits performed over windows dominated by dijet events are typically of a higher quality than those that extend into the far tails of the distributions, possibly motivating future fits focused more heavily in this region. Finally, we discuss how different estimates of the three-loop soft matching coefficient c S ˜ 3 can also lead to measurable changes in the fitted { α s , Ω 1 } values. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Revealing the evolution of order in materials microstructures using multi-modal computer vision

The development of high-performance materials for microelectronics, energy storage, and extreme environments depends on our ability to describe and direct property-defining microstructural order. Our present understanding is typically derived from laborious manual analysis of imaging and spectroscopy data, which is difficult to scale, challenging to reproduce, and lacks the ability to reveal latent associations needed for mechanistic models. Here, we demonstrate a multi-modal machine learning (ML) approach to describe order from electron microscopy analysis of the complex oxide La 1−x Sr x FeO 3 . We construct a hybrid pipeline based on fully and semi-supervised classification, allowing us to evaluate both the characteristics of each data modality and the value each modality adds to the ensemble. We observe distinct differences in the performance of uni- and multi-modal models, from which we draw general lessons in describing crystal order using computer vision.

36 MATERIALS SCIENCE↗

Three-dimensional modeling of hyphal fusion, branching, and nutrient transport in filamentous fungi

Fungi exhibit behaviors distinct from other microbes. Filamentous fungi grow by extending complex networks of branched filaments collectively referred to as the mycelium. These networks can expand over large distances and traverse low-nutrient areas by translocating nutrients through the filament network. This spatial characteristic makes filamentous fungi crucial for soil ecosystems, supporting stable microbial communities and promoting plant growth. However, simulating these behaviors is complex. The elongated nature of fungal compartments results in different mechanical interactions compared to the commonly modeled spherical bacteria. These detailed hyphal mechanics require specialized consideration and are often excluded from conventional fungal simulation packages. Additionally, the extensive fungal networks in nature demand computationally intensive simulations, necessitating high-performance algorithms. Therefore, realistic fungi simulations require specialized software. Here, we introduce a fungal modeling expansion to the high-performance biological modelling and interface exchange (bmx) software suite. bmx leverages adaptive mesh refinement in AMReX for chemical diffusion and incorporates a full mechanical model for bacterial cells, accelerated by GPUs. By extending bmx to model filamentous particles, we demonstrate the formation of complex filament networks through interactions like hyphal branching and fusion (anastomosis). We show that the networks produced match real-world fungal structures through various metrics. This work supports computational studies of fungal growth dynamics and can be adapted to investigate the growth of other filamentous structures in biology or materials science. The expanded-BMX package is open-sourced and is available online.

Cell mechanics↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗

Biased degenerate ground-state sampling of small Ising models with converged quantum approximate optimization algorithm

The quantum alternating operator ansatz, a generalization of the quantum approximate optimization algorithm (QAOA), is a quantum algorithm used for approximately solving combinatorial optimization problems. QAOA typically uses the transverse field mixer as the driving Hamiltonian. One of the interesting properties of the transverse field driving Hamiltonian is that it results in nonuniform sampling of degenerate ground states of optimization problems. In this study, we numerically examine the fair sampling properties of the transverse field mixer QAOA, and Grover mixer QAOA (GM-QAOA), which provides theoretical guarantees of fair sampling of degenerate optimal solutions, up to a large enough p such that the mean expectation value converges to an optimal approximation ratio of 1. This comparison is performed with high-quality heuristically computed, but not necessarily optimal, QAOA angles, which give strictly monotonically improving solution quality as p increases. These angles are computed using the Julia based numerical simulation software JuliQAOA. Fair sampling of degenerate ground states is quantified using the Shannon entropy of the ground-state amplitudes distribution. The fair sampling properties are reported on several quantum signature Hamiltonians from previous quantum annealing fair sampling studies. Small random fully connected spin glasses are shown, which exhibit exponential suppression of some degenerate ground states with transverse field mixer QAOA. The transverse field mixer QAOA simulations show that some problem instances clearly saturate the Shannon entropy of 0 with a maximally biased distribution that occurs when the learning converges to an approximation ratio of 1 while other problem instances never deviate from a maximum Shannon entropy (uniform distribution) at any p step. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING↗

Data mining and computational screening of Rashba-Dresselhaus splitting and optoelectronic properties in two-dimensional perovskite materials

Recent developments highlighting the promise of two-dimensional perovskites have vastly increased the compositional search space in the perovskite family. This presents a great opportunity for the realization of highly performant devices and practical challenges associated with the identification of candidate materials. High-fidelity computational screening offers great value in this regard. In this study, we carry out a multiscale computational workflow, generating a dataset of two-dimensional perovskites in the Dion-Jacobson and Ruddlesden-Popper phases. Our dataset comprises ten B-site cations, four halogens, and over 20 organic cations across over 2000 materials. We compute electronic properties, thermoelectric performance, and numerous geometric characteristics. Furthermore, we introduce a framework for the high-throughput computation of Rashba-Dresselhaus splitting. Finally, we use this dataset to train machine learning models for the accurate prediction of band gaps, candidate Rashba-Dresselhaus materials, and partial charges. The work presented herein can aid future investigations of two-dimensional perovskites with targeted applications in mind.

14 SOLAR ENERGY↗

Raptor

Raptor is an efficient Python-based tool for predicting the formation and morphology of stochastic lack of fusion defects in metal AM processes. A major obstacle for the qualification and certification of additively manufactured parts in critical applications continues to be performance variability caused in part by porosity-related defects. High-fidelity process models that could predict these defect features are currently too computationally expensive for component-level analysis. To address this, Raptor employs a high-performance geometric method to model the dynamic melt pool rather than relying on computationally intensive thermal fluid dynamics. This allows Raptor to rapidly identify regions of unmelted material that correspond to lack of fusion pores. The efficiency of this approach significantly reduces the time and resources needed for generating 3D defect predictions, which enables users to conduct large-scale parameter studies and evaluate how process variations affect part quality. The framework offers operational flexibility; users can execute simulations through a simple command line interface or integrate core functions as a library within larger computational workflows. Simulation outputs include 3D porosity maps for visualization and tools for quantitative morphological analysis. These results are suitable for direct comparison with experimental characterization data from methods such as X-ray computed tomography and can be used for statistical process optimization.

Subraveti, Vamsi [Vanderbilt Univ., Nashville, TN ↗

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (↗

SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical Segmentation

This paper presents a novel approach to addressing the long-sequence problem in high-resolution medical images for Vision Transformers (ViTs). Using smaller patches as tokens can enhance ViT performance, but quadratically increases computation and memory requirements. Therefore, the common practice for applying ViTs to high-resolution images is either to: (a) employ complex sub-quadratic attention schemes or (b) use large to medium-sized patches and rely on additional mechanisms within the model to capture the spatial hierarchy of details. We propose Symmetrical Hierarchical Forest (SHF), a lightweight approach that adaptively patches the input image to increase token information density and encode hierarchical spatial structures into the input embedding. We then apply a reverse depatching scheme to the output embeddings of the transformer encoder, eliminating the need for convolution-based decoders. Unlike previous methods that modify attention mechanisms or use a complex hierarchy of interacting models, SHF can be retrofitted to any ViT model to allow it to learn the hierarchical structure of details in high-resolution images without requiring architectural changes. Experimental results demonstrate significant gains in computational efficiency and performance: on the PAIP WSI dataset, we achieved a 3∼32×speedup or a 2.95%∼7.03% increase in accuracy (measured by Dice score) at a 64K2 resolution with the same computational budget, compared to state-of-the-art production models. On the 3D medical datasets BTCV and KiTS, training was 6×faster, with accuracy gains of 6.93% and 5.9%, respectively, compared to models without SHF.

Zhang, Enzhi [Hokkaido University, Japan]↗

Synaptic Functionality and Neuromorphic Information Processing in Membrane Ion Channel Junctions

The human brain performs complex memory and computational tasks with high energy efficiency by regulating ion transport through membrane channels. These signaling mechanisms have been inspiring the development of nanofluidic memristors that emulate synaptic behavior. Here, in this study, we describe a membrane ion channel synapse (MICS), constructed from aqueous droplets linked by gramicidin A channels, that achieves neuromorphic functionality. MICS exhibits memristive ion transport with hysteretic current–voltage behavior arising from voltage-dependent channel formation and ion transport dynamics. MICS emulates a range of synaptic behaviors including associative learning. We further demonstrate its application in reservoir computing by performing handwritten digit classification and tic-tac-toe game and explore the system parameters that improve the computational performance. This droplet-based biomimetic synapse offers a potentially scalable and energy-efficient platform for next-generation neuromorphic computing systems.

Droplet interface bilayer↗

Accelerating computing for the future electric grid (CRADA Final Report)

As a participant in the Cyclotron Road Lab-Embedded Entrepreneurship Program (LEEP), Vellex Computing, Inc. has successfully validated the "Vellex Computing Stack," a breakthrough Analog Neural Computer (ANC) specifically designed for high-performance edge optimization. This project achieved critical milestones in mixed-signal circuit stability and software-hardware co-design, directly addressing national priorities in semiconductor resiliency. The success of this work is deeply rooted in the support from the Cyclotron Road LEEP, which provided the essential "hard tech" runway—funding, mentorship, and access to Lawrence Berkeley National Laboratory’s world-class characterization facilities—allowing Vellex to overcome the "Valley of Death" often faced by deep-tech hardware startups. By leveraging LBNL’s advanced testing infrastructure, Vellex was able to rigorously benchmark the ANC architecture against state-of-the-art digital solutions, a feat that would have been resource-prohibitive independently. This collaboration has not only advanced American leadership in analog computing but has also matured Vellex’s technology to a stage ripe for private sector commercialization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A substitutional quantum defect in WS2 discovered by high-throughput computational screening and fabricated by site-selective STM manipulation

Abstract Point defects in two-dimensional materials are of key interest for quantum information science. However, the parameter space of possible defects is immense, making the identification of high-performance quantum defects very challenging. Here, we perform high-throughput (HT) first-principles computational screening to search for promising quantum defects within WS 2 , which present localized levels in the band gap that can lead to bright optical transitions in the visible or telecom regime. Our computed database spans more than 700 charged defects formed through substitution on the tungsten or sulfur site. We found that sulfur substitutions enable the most promising quantum defects. We computationally identify the neutral cobalt substitution to sulfur ( $${\rm{Co}}_{{{{{{{{\rm{S}}}}}}}}}^{0}$$ Co S 0 ) and fabricate it with scanning tunneling microscopy (STM). The $${\rm{Co}}_{{{{{{{{\rm{S}}}}}}}}}^{0}$$ Co S 0 electronic structure measured by STM agrees with first principles and showcases an attractive quantum defect. Our work shows how HT computational screening and nanoscale synthesis routes can be combined to design promising quantum defects.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Mitigating the Effects of Au-Al Intermetallic Compounds Due to High-Temperature Processing of Surface-Electrode Ion Traps

Stringent physical requirements need to be met for the high-performing surface-electrode ion traps used in quantum computing and timekeeping. In particular, these traps must survive a high-temperature environment for vacuum chamber preparation and support high RF voltage on closely spaced electrodes. Due to the use of gold wire bonds on aluminum pads, intermetallic growth can lead to wire bond failure via breakage or high resistance, limiting the lifetime of a trap assembly to a single multiday bake at 200 ° C. Using traditional thick metal stacks to prevent intermetallic growth, however, can result in trap failure due to RF breakdown events. Through high-temperature experiments, we conclude that an ideal metal stack for ion traps is Ti/Pt/Au (20/100/250 nm), which allows for a cumulative bakeable time of roughly 86 days without compromising the trap voltage performance. This increase in the bakeable lifetime of ion traps will remove the need to discard otherwise functional ion traps when vacuum hardware is upgraded, which will greatly benefit ion trap experiments.

Haltli, Raymond A.↗

High spatial resolution neutron imaging of lithium-ion batteries: Correlating microstructure and lithium transport

Thick electrodes for lithium-ion batteries can increase the overall energy density, but increasing the electrode thickness introduces charge transport limitations. These limitations may be mitigated through proper electrode structuring. Here, high spatial resolution neutron imaging was used to understand the correlation between microstructure and lithium transport in lithium-ion anodes. Batteries with distinct graphite anode microstructures were produced and studied with high spatial resolution in operando neutron radiography to observe the effects of structure on transport. High spatial resolution neutron computed tomography was performed following in operando neutron radiography. X-ray computed tomography and scanning electron microscopy were used to observe the finer scale anode structure to complement neutron imaging. Solvent-free anodes containing a tightly-packed layered structure confined lithium movement close to the separator. This structure limited capacity, but supported better rate capability. Conversely, a more open pore structure in the wet cast anodes yielded higher capacity with reduced rate capability. Together, these results show that lithium distributions can be controlled by the macroscopic structure of the electrodes, the microstructural pore network, and the microscale active areas that support electrochemical reactions. Furthermore, multimodal imaging applying the complementary strengths of neutron and X-ray methods is shown as a tool for advancing battery design.

25 ENERGY STORAGE↗

Bringing randomized algorithms to mainstream numerical linear algebra

Numerical linear algebra (NLA) underpins huge swaths of computational science and engineering. For scientists and engineers to make the most of the DOE’s computing resources, it is essential that they have access to high-performance implementations of algorithms with best-in-class scalability and reliability. Despite this, prevailing NLA libraries have little to no support for breakthrough algorithms from the field of randomized numerical linear algebra (RandNLA) that have been developed over the past twenty years. The goal of this LDRD was to break a log-jam that had prevented broad adoption of RandNLA. Our work had two thrusts. The first was to develop RandBLAS: a trustworthy and high-performance C++ library for randomized dimension reduction (an operation widely known as sketching). The second was the development of a novel randomized algorithm for computing a challenging type of matrix decomposition known as Householder QR with column pivoting (Householder QRCP). In this one-year late-start LDRD we successfully delivered RandBLAS 1.0 and new CPU and GPU codes for Householder QRCP. RandBLAS has extensive documentation at https://randblas.readthedocs.io/en/stable/. Papers on RandBLAS and and our high-performance QRCP codes are forthcoming.

97 MATHEMATICS AND COMPUTING↗