Search NASA⌕ Search

SEARCH · Search NASA

Results for “near memory computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A survey on processing-in-memory techniques: Advances and challenges

Processing-in-memory (PIM) techniques have gained much attention from computer architecture researchers, and significant research effort has been invested in exploring and developing such techniques. Increasing the research activity dedicated to improving PIM techniques will hopefully help deliver PIM’s promise to solve or significantly reduce memory access bottleneck problems for memory-intensive applications. We also believe it is imperative to track the advances made in PIM research to identify open challenges and enable the research community to make informed decisions and adjust future research directions. In this survey, we analyze recent studies that explored PIM techniques, summarize the advances made, compare recent PIM architectures, and identify target application domains and suitable memory technologies. We also discuss proposals that address unresolved issues of PIM designs (e.g., address translation/mapping of operands, workload analysis to identify application segments that can be accelerated with PIM, OS/runtime support, and coherency issues that must be resolved to incorporate PIM). We believe this work can serve as a useful reference for researchers exploring PIM techniques.

97 MATHEMATICS AND COMPUTING↗

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (↗

Parallel Simulation of Unsteady Turbulent Flames

Time-accurate simulation of turbulent flames in high Reynolds number flows is a challenging task since both fluid dynamics and combustion must be modeled accurately. To numerically simulate this phenomenon, very large computer resources (both time and memory) are required. Although current vector supercomputers are capable of providing adequate resources for simulations of this nature, the high cost and their limited availability, makes practical use of such machines less than satisfactory. At the same time, the explicit time integration algorithms used in unsteady flow simulations often possess a very high degree of parallelism, making them very amenable to efficient implementation on large-scale parallel computers. Under these circumstances, distributed memory parallel computers offer an excellent near-term solution for greatly increased computational speed and memory, at a cost that may render the unsteady simulations of the type discussed above more feasible and affordable.This paper discusses the study of unsteady turbulent flames using a simulation algorithm that is capable of retaining high parallel efficiency on distributed memory parallel architectures. Numerical studies are carried out using large-eddy simulation (LES). In LES, the scales larger than the grid are computed using a time- and space-accurate scheme, while the unresolved small scales are modeled using eddy viscosity based subgrid models. This is acceptable for the moment/energy closure since the small scales primarily provide a dissipative mechanism for the energy transferred from the large scales. However, for combustion to occur, the species must first undergo mixing at the small scales and then come into molecular contact. Therefore, global models cannot be used. Recently, a new model for turbulent combustion was developed, in which the combustion is modeled, within the subgrid (small-scales) using a methodology that simulates the mixing and the molecular transport and the chemical kinetics within each LES grid cell. Finite-rate kinetics can be included without any closure and this approach actually provides a means to predict the turbulent rates and the turbulent flame speed. The subgrid combustion model requires resolution of the local time scales associated with small-scale mixing, molecular diffusion and chemical kinetics and, therefore, within each grid cell, a significant amount of computations must be carried out before the large-scale (LES resolved) effects are incorporated. Therefore, this approach is uniquely suited for parallel processing and has been implemented on various systems such as: Intel Paragon, IBM SP-2, Cray T3D and SGI Power Challenge (PC) using the system independent Message Passing Interface (MPI) compiler. In this paper, timing data on these machines is reported along with some characteristic results.

Menon, Suresh↗

Nitrogen Vacancies Induce Fatigue in Ferroelectric Al 0.93 B 0.07 N

Wurtzite ferroelectrics (e.g., Al 0.93 B 0.07 N) are being explored for high-temperature and emerging near-, or in-compute, memory architectures due to the material advantages offered by their large remanent polarization and robust chemical stability. Despite these advantages, current Al0.93B0.07N devices do not have sufficient endurance lifetime to meet roadmap targets. To identify the defects responsible for this limited endurance, a combination of electronic measurements and optical spectroscopies characterized the evolution of defect states within Al 0.93 B 0.07 N with cycling. Ultrathin (∼10 nm) metal contacts were used to optically probe regions subject to ferroelectric switching; photoluminescence spectroscopy identified the emergence of a transition near 2.1 eV whose intensity scaled with the non-switching polarization quantified via positive-up negative-down (PUND) measurements. Accompanying thermally stimulated depolarization current (TSDC) and modulus spectroscopy measurements also observed the strengthening of a state near 2.1 eV. The origin of this feature is ascribed to transitions between a nitrogen vacancy and another defect deeper in the bandgap. Recognizing that the impurity concentration is largely fixed, strengthening of this transition indicates an increase in the number of nitrogen vacancies. Switching, therefore, creates vacancies in Al 0.93 B 0.07 N likely due to hot-atom damage induced by the aggressive fields necessary to switch wurtzite materials that ultimately limits endurance.

36 MATERIALS SCIENCE↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Simulation of Laser Cooling and Trapping in Engineering Applications

An advanced computer code is undergoing development for numerically simulating laser cooling and trapping of large numbers of atoms. The code is expected to be useful in practical engineering applications and to contribute to understanding of the roles that light, atomic collisions, background pressure, and numbers of particles play in experiments using laser-cooled and -trapped atoms. The code is based on semiclassical theories of the forces exerted on atoms by magnetic and optical fields. Whereas computer codes developed previously for the same purpose account for only a few physical mechanisms, this code incorporates many more physical mechanisms (including atomic collisions, sub-Doppler cooling mechanisms, Stark and Zeeman energy shifts, gravitation, and evanescent-wave phenomena) that affect laser-matter interactions and the cooling of atoms to submillikelvin temperatures. Moreover, whereas the prior codes can simulate the interactions of at most a few atoms with a resonant light field, the number of atoms that can be included in a simulation by the present code is limited only by computer memory. Hence, the present code represents more nearly completely the complex physics involved when using laser-cooled and -trapped atoms in engineering applications. Another advantage that the code incorporates is the possibility to analyze the interaction between cold atoms of different atomic number. Some properties that cold atoms of different atomic species have, like cross sections and the particular excited states they can occupy when interacting with each other and light fields, play important roles not yet completely understood in the new experiments that are under way in laboratories worldwide to form ultracold molecules. Other research efforts use cold atoms as holders of quantum information, and more recent developments in cavity quantum electrodynamics also use ultracold atoms to explore and expand new information-technology ideas. These experiments give a hint on the wide range of applications and technology developments that can be tackled using cold atoms and light fields. From more precise atomic clocks and gravity sensors to the development of quantum computers, there will be a need to completely understand the whole ensemble of physical mechanisms that play a role in the development of such technologies. The code also permits the study of the dynamic and steady-state operations of technologies that use cold atoms. The physical characteristics of lasers and fields can be time-controlled to give a realistic simulation of the processes involved such that the design process can determine the best control features to use. It is expected that with the features incorporated into the code it will become a tool for the useful application of ultracold atoms in engineering applications. Currently, the software is being used for the analysis and understanding of simple experiments using cold atoms, and for the design of a modular compact source of cold atoms to be used in future research and development projects. The results so far indicate that the code is a useful design instrument that shows good agreement with experimental measurements (see figure), and a Windows-based user-friendly interface is also under development.

Ramirez-Serrano, Jaime↗

Portable parallel stochastic optimization for the design of aeropropulsion components

This report presents the results of Phase 1 research to develop a methodology for performing large-scale Multi-disciplinary Stochastic Optimization (MSO) for the design of aerospace systems ranging from aeropropulsion components to complete aircraft configurations. The current research recognizes that such design optimization problems are computationally expensive, and require the use of either massively parallel or multiple-processor computers. The methodology also recognizes that many operational and performance parameters are uncertain, and that uncertainty must be considered explicitly to achieve optimum performance and cost. The objective of this Phase 1 research was to initialize the development of an MSO methodology that is portable to a wide variety of hardware platforms, while achieving efficient, large-scale parallelism when multiple processors are available. The first effort in the project was a literature review of available computer hardware, as well as review of portable, parallel programming environments. The first effort was to implement the MSO methodology for a problem using the portable parallel programming language, Parallel Virtual Machine (PVM). The third and final effort was to demonstrate the example on a variety of computers, including a distributed-memory multiprocessor, a distributed-memory network of workstations, and a single-processor workstation. Results indicate the MSO methodology can be well-applied towards large-scale aerospace design problems. Nearly perfect linear speedup was demonstrated for computation of optimization sensitivity coefficients on both a 128-node distributed-memory multiprocessor (the Intel iPSC/860) and a network of workstations (speedups of almost 19 times achieved for 20 workstations). Very high parallel efficiencies (75 percent for 31 processors and 60 percent for 50 processors) were also achieved for computation of aerodynamic influence coefficients on the Intel. Finally, the multi-level parallelization strategy that will be needed for large-scale MSO problems was demonstrated to be highly efficient. The same parallel code instructions were used on both platforms, demonstrating portability. There are many applications for which MSO can be applied, including NASA's High-Speed-Civil Transport, and advanced propulsion systems. The use of MSO will reduce design and development time and testing costs dramatically.

Sues, Robert H.↗

An image compression algorithm for a high-resolution digital still camera

The Electronic Still Camera (ESC) project will provide for the capture and transmission of high-quality images without the use of film. The image quality will be superior to video and will approach the quality of 35mm film. The camera, which will have the same general shape and handling as a 35mm camera, will be able to send images to earth in near real-time. Images will be stored in computer memory (RAM) in removable cartridges readable by a computer. To save storage space, the image will be compressed and reconstructed at the time of viewing. Both lossless and loss-y image compression algorithms are studied, described, and compared.

Nerheim, Rosalee↗

On the Rapid Computation of Various Polylogarithmic Constants

We give algorithms for the computation of the d-th digit of certain transcendental numbers in various bases. These algorithms can be easily implemented (multiple precision arithmetic is not needed), require virtually no memory, and feature run times that scale nearly linearly with the order of the digit desired. They make it feasible to compute, for example, the billionth binary digit of log(2) or pi on a modest workstation in a few hours run time. We demonstrate this technique by computing the ten billionth hexadecimal digit of pi, the billionth hexadecimal digits of pi-squared, log(2) and log-squared(2), and the ten billionth decimal digit of log(9/10). These calculations rest on the observation that very special types of identities exist for certain numbers like pi, pi-squared, log(2) and log-squared(2). These are essentially polylogarithmic ladders in an integer base. A number of these identities that we derive in this work appear to be new, for example a critical identity for pi.

Bailey, David H.↗

ALD-Derived WO 3– x Leads to Nearly Wake-Up-Free Ferroelectric Hf 0.5 Zr 0.5 O 2 at Elevated Temperatures

Breaking the memory wall in advanced computing architectures will require complex 3D integration of emerging memory materials such as ferroelectrics─either within the back-end-of-line (BEOL) of CMOS front-end processes or through advanced 3D packaging technologies. Achieving this integration demands that memory materials exhibit high thermal resilience, with the capability to operate reliably at elevated temperatures, such as 125°C, due to the substantial heat generated by front-end transistors. However, silicon-compatible HfO 2 -based ferroelectrics tend to exhibit antiferroelectric-like behavior in this temperature range, accompanied by a more pronounced wake-up effect, posing significant challenges to their thermal reliability. Here, we report that by introducing a thin tungsten oxide (WO 3–x ) layer─known as an oxygen reservoir─and carefully tuning its oxygen content, ultrathin Hf 0.5 Zr 0.5 O 2 (5 nm) films can be made robust against the ferroelectric-to-antiferroelectric transition at elevated temperatures. This approach not only minimizes polarization loss in the pristine state but also effectively suppresses the wake-up effect, reducing the required wake-up cycles from 10 5 to only 10 at 125°C, a qualifying temperature for back-end memory integrated with front-end logic, as defined by the JEDEC standard. First-principles density functional theory (DFT) calculations reveal that WO 3 enhances the stability of the ferroelectric orthorhombic phase (o-phase) at elevated temperatures by increasing the tetragonal-to-orthorhombic phase energy gap and promoting favorable phonon mode evolution, thereby supporting o-phase formation under both thermodynamic and kinetic constraints.

36 MATERIALS SCIENCE↗

Polaron-induced metal-to-insulator transition in vanadium oxides from density functional theory calculations

Vanadium oxides have been extensively studied as phase-change memory units in artificial synapses for neuromorphic computing due to their metal-insulator transitions (MIT) at or near room temperature. Recently, injection of charge carriers into vanadium oxides, e.g., via optically via a heterostructure, has been proposed as an alternative switching mechanism and also potentially as a means to tune the MIT temperature. In this study, we explore the formation of small polarons in the low temperature (LT) insulating phases for V 3 O 5 ,VO 2 , and V 2 O 3 , and the barriers to their migration using density functional theory calculations. We find that V 3 O 5 exhibits very low hole and electron polaron migration barriers (<100 meV) compared to V 2 O 3 and VO 2 , leading to much higher estimated polaronic conductivity. We also link the relative migration barriers to the amount of distortion that has to travel when the polaron migrate from one site to another. Polarons in V 3 O 5 also have smaller binding energies to vanadium and oxygen vacancy defects. Furthermore, these results suggest that the triggering of the MIT via injection of charge carriers are due to the formation of small polarons that can migrate rapidly through the crystal.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Efficient learning of t -doped stabilizer states with single-copy measurements

One of the primary objectives in the field of quantum state learning is to develop algorithms that are time-efficient for learning states generated from quantum circuits. Earlier investigations have demonstrated time-efficient algorithms for states generated from Clifford circuits with at most log &#x2061; ( n ) non-Clifford gates. However, these algorithms necessitate multi-copy measurements, posing implementation challenges in the near term due to the requisite quantum memory. On the contrary, using solely single-qubit measurements in the computational basis is insufficient in learning even the output distribution of a Clifford circuit with one additional T gate under reasonable post-quantum cryptographic assumptions. In this work, we introduce an efficient quantum algorithm that employs only nonadaptive single-copy measurement to learn states produced by Clifford circuits with a maximum of O ( log &#x2061; n ) non-Clifford gates, filling a gap between the previous positive and negative results.

Physics↗

A high-performance FFT algorithm for vector supercomputers

Many traditional algorithms for computing the fast Fourier transform (FFT) on conventional computers are unacceptable for advanced vector and parallel computers because they involve nonunit, power-of-two memory strides. A practical technique for computing the FFT that avoids all such strides and appears to be near-optimal for a variety of current vector and parallel computers is presented. Performance results of a program based on this technique are given. Notable among these results is that a FORTRAN implementation of this algorithm on the CRAY-2 runs up to 77-percent faster than Cray's assembly-coded library routine.

Bailey, David H.↗

Comparison of uniform perturbation solutions and numerical solutions for some potential flows past slender bodies

Approximate solutions for potential flow past an axisymmetric slender body and past a thin airfoil are calculated using a uniform perturbation method and then compared with either the exact analytical solution or the solution obtained using a purely numerical method. The perturbation method is based upon a representation of the disturbance flow as the superposition of singularities distributed entirely within the body, while the numerical (panel) method is based upon a distribution of singularities on the surface of the body. It is found that the perturbation method provides very good results for small values of the slenderness ratio and for small angles of attack. Moreover, for comparable accuracy, the perturbation method is simpler to implement, requires less computer memory, and generally uses less computation time than the panel method. In particular, the uniform perturbation method yields good resolution near the regions of the leading and trailing edges where other methods fail or require special attention.

Wong, T. C.↗

Comparison of uniform perturbation and numerical solutions for some potential flows past slender bodies

Approximate solutions for potential flow past an axisymmetric slender body and past a thin airfoil are calculated using a uniform perturbation method and then compared with either the exact analytical solution or the solution obtained using a purely numerical method. The perturbation method is based upon a representation of the disturbance flow as the superposition of singularities distributed entirely within the body, while the numerical (panel) method is based upon a distribution of singularities on the surface of the body. It is found that the perturbation method provides very good results for small values of the slenderness ratio and for small angles of attack. Moreover, for comparable accuracy, the perturbation method is simpler to implement, requires less computer memory, and generally uses less computation time than the panel method. In particular, the uniform perturbation method yields good resolution near the regions of the leading and trailing edges where other methods fail or require special attention.

Wong, T.-C.↗

Modeling Materials: Design for Planetary Entry, Electric Aircraft, and Beyond

NASA missions push the limits of what is possible. The development of high-performance materials must keep pace with the agency's demanding, cutting-edge applications. Researchers at NASA's Ames Research Center are performing multiscale computational modeling to accelerate development times and further the design of next-generation aerospace materials. Multiscale modeling combines several computationally intensive techniques ranging from the atomic level to the macroscale, passing output from one level as input to the next level. These methods are applicable to a wide variety of materials systems. For example: (a) Ultra-high-temperature ceramics for hypersonic aircraft-we utilized the full range of multiscale modeling to characterize thermal protection materials for faster, safer air- and spacecraft, (b) Planetary entry heat shields for space vehicles-we computed thermal and mechanical properties of ablative composites by combining several methods, from atomistic simulations to macroscale computations, (c) Advanced batteries for electric aircraft-we performed large-scale molecular dynamics simulations of advanced electrolytes for ultra-high-energy capacity batteries to enable long-distance electric aircraft service; and (d) Shape-memory alloys for high-efficiency aircraft-we used high-fidelity electronic structure calculations to determine phase diagrams in shape-memory transformations. Advances in high-performance computing have been critical to the development of multiscale materials modeling. We used nearly one million processor hours on NASA's Pleiades supercomputer to characterize electrolytes with a fidelity that would be otherwise impossible. For this and other projects, Pleiades enables us to push the physics and accuracy of our calculations to new levels.

Supercomputing↗

Parallel spatial direct numerical simulations on the Intel iPSC/860 hypercube

The implementation and performance of a parallel spatial direct numerical simulation (PSDNS) approach on the Intel iPSC/860 hypercube is documented. The direct numerical simulation approach is used to compute spatially evolving disturbances associated with the laminar-to-turbulent transition in boundary-layer flows. The feasibility of using the PSDNS on the hypercube to perform transition studies is examined. The results indicate that the direct numerical simulation approach can effectively be parallelized on a distributed-memory parallel machine. By increasing the number of processors nearly ideal linear speedups are achieved with nonoptimized routines; slower than linear speedups are achieved with optimized (machine dependent library) routines. This slower than linear speedup results because the Fast Fourier Transform (FFT) routine dominates the computational cost and because the routine indicates less than ideal speedups. However with the machine-dependent routines the total computational cost decreases by a factor of 4 to 5 compared with standard FORTRAN routines. The computational cost increases linearly with spanwise wall-normal and streamwise grid refinements. The hypercube with 32 processors was estimated to require approximately twice the amount of Cray supercomputer single processor time to complete a comparable simulation; however it is estimated that a subgrid-scale model which reduces the required number of grid points and becomes a large-eddy simulation (PSLES) would reduce the computational cost and memory requirements by a factor of 10 over the PSDNS. This PSLES implementation would enable transition simulations on the hypercube at a reasonable computational cost.

Joslin, Ronald D.↗

Equation solvers for distributed-memory computers

A large number of scientific and engineering problems require the rapid solution of large systems of simultaneous equations. The performance of parallel computers in this area now dwarfs traditional vector computers by nearly an order of magnitude. This talk describes the major issues involved in parallel equation solvers with particular emphasis on the Intel Paragon, IBM SP-1 and SP-2 processors.

Storaasli, Olaf O.↗