Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Active interlocking metasurfaces enabled by shape memory alloys

Interlocking metasurfaces (ILMs) are a newly developed joining technology that relies on arrays of interlocking features that transmit force and constrain motion between adjoining bodies in one or more directions. This study explores harnessing the shape memory effect (SME) in Nickel-Titanium shape memory alloys (NiTi SMAs) in structures fabricated using additive manufacturing (AM) to advance the development of active ILMs by creating unit cells that open or close at specific temperatures. The study encompasses designing and fabricating two distinct interlocking array configurations using near-equiatomic NiTi powder and the laser powder bed fusion (L-PBF) AM technique, following a previously developed AM process optimization framework to manufacture defect-free parts. To guide the design process, finite element analysis (FEA) was employed to predict strain values during engage-disengage cycles. The martensitic transformation characteristics of the ILMs were characterized. Thermomechanical testing revealed that the ILMs demonstrate high locking force once engaged, coupled with complete shape recovery and good cyclic stability. Digital image correlation (DIC) was also employed to validate the FEA predictions during the engage-disengage cycles. The results indicate that NiTi SMA-based ILMs can be designed and fabricated into complex shapes using L-PBF. By leveraging the SME, the functionality of an ILM can be improved upon. The combination of computational modeling, additive manufacturing, and thermomechanical and physical property characterization provides a framework for designing future ILMs out of active materials.

Additive manufacturing↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Ultracoherent SRF Cavity-Based Multi-Qudit Platform with Error-Resilient Control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6\,ms and 15.6\,ms for the two modes, and a pure dephasing time exceeding 40\,ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to $N = 20$ with fidelities exceeding 95\%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9\% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Kim, T. [Northwestern U.]↗

Ultracoherent superconducting cavity-based multiqudit platform with error-resilient control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical-shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6 ms and 15.6 ms for the two modes, and a pure dephasing time exceeding 40 ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to N = 20 with fidelities exceeding 95%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Lu, Yao [Fermilab] (ORCID:000000020413698X)↗

Ultracoherent superconducting cavity-based multiqudit platform with error-resilient control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical-shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6 ms and 15.6 ms for the two modes, and a pure dephasing time exceeding 40 ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to N = 20 with fidelities exceeding 95%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Lu, Yao [Fermilab] (ORCID:000000020413698X)↗

Ultracoherent superconducting cavity-based multiqudit platform with error-resilient control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6 ms and 15.6 ms for the two modes, and a pure dephasing time exceeding 40 ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to $N = 20$ with fidelities exceeding 95%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Kim, Taeyoon [Fermilab; Northwestern U.]↗

PETSc/TAO Users Manual Revision 3.22

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.23

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.24

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual Revision 3.25

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved Evaluation of Large Network Matrices for Linear Power Flow Within Optimization Problems

This work presents methods for evaluating the Power Transfer Distribution Factor (PTDF) and Line Outage Distribution Factor (LODF) matrices by employing sparse linear algebra for large-scale computing applications. These matrices play a critical role in many power system applications, such as the Unit Commitment Problem (UC), pre- and post-contingency power flow analysis, and transmission expansion. These matrices are typically dense, which means they require a significant amount of time and memory to be computed for large networks. However, by analyzing the structure of the matrices and their computation method, it is possible to use reduced memory methods based on sparse matrix operations. This paper shows that sparse linear algebra algorithms are faster and require less memory and time than traditional dense approaches. Additionally, we explore the effect of matrix sparsification by eliminating trailing digits on power flow calculations.

large scale↗

kokkos-fft: A shared-memory FFT for the Kokkos ecosystem

kokkos-fft provides a unified, performance-portable interface for Fast Fourier Transforms (FFTs) within the Kokkos ecosystem (C. Trott et al., 2021). It seamlessly integrates with leading local FFT libraries including FFTW, cuFFT, rocFFT, and oneMKL. Designed for simplicity and efficiency, kokkos-fft offers a user experience akin to numpy.fft for in-place and out-of-place transforms, while leveraging the raw speed of vendor-optimized libraries. A demonstration solving 2D Hasegawa-Wakatani turbulence with the Fourier spectral method illustrates how kokkos-fft can deliver significant speedups over Python-based alternatives without drastically increasing code complexity, empowering researchers to perform high-performance FFTs simply and effectively.

97 MATHEMATICS AND COMPUTING↗

Compact fiber-coupled narrowband two-mode squeezed light source

Quantum correlated states of light, such as squeezed states, are a fundamental resource for the development of quantum technologies, as they are needed for applications in quantum metrology, quantum computation, and quantum communications. It is thus critical to develop compact, efficient, and robust sources to generate such states. Here, we report on a compact, narrowband, fiber-coupled source of two-mode squeezed states of light at 795 nm based on four-wave mixing (FWM) in an 85 Rb atomic vapor. The source is designed in a small modular form factor, with two input fiber-coupled beams, the seed and pump beams required for the FWM, and two output fibers, one for each of the modes of the squeezed state. The system is optimized for low pump power (135 mW) to achieve a maximum intensity-difference squeezing of 4.4 dB after the output of fibers at an analysis frequency of 1 MHz. Furthermore, the narrowband nature of the source makes it ideal for atomic-based quantum sensing and quantum networking configurations that rely on atomic quantum memories. Such a source paves the way for a versatile and portable platform for applications in quantum information science.

Jain, Umang [University of Oklahoma, Norman, OK (U↗

Properties of Electronic Materials

This final technical report summarizes the research conducted under DOE Grant DE-SC0002623, "Properties of Electronic Materials," led by Principal Investigator Shengbai Zhang at Rensselaer Polytechnic Institute. Over the 16-year period, the project employed first-principles computational methods to investigate the structural, electronic, and dynamic properties of a wide range of electronic materials, with applications in energy technologies, optoelectronics, and data storage. Key areas included topological insulators, phase-change materials, graphene and two-dimensional systems, perovskites for photovoltaics, defect engineering in semiconductors, kagome lattices, and ultrafast carrier dynamics. The research resulted in 115 peer-reviewed publications, advancing fundamental understanding of material behaviors at the atomic scale and contributing to innovations in renewable energy, memory devices, and quantum materials. Findings have implications for improving energy efficiency, developing lead-free solar cells, and enabling high-speed data processing. The work has trained numerous graduate students and postdocs, fostering the next generation of computational materials scientists. The original goals were to develop theoretical models and computational tools to predict and optimize electronic properties of materials for energy applications. All objectives were accomplished, with no major departures from planned methodologies. Challenges in computational scaling were addressed through access to high-performance computing resources.

36 MATERIALS SCIENCE↗

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗