Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

Fourier-MIONet: Fourier-enhanced multiple-input neural operators for multiphase modeling of geological carbon sequestration

Geologic carbon sequestration (GCS) is a safety-critical technology that aims to reduce the amount of carbon dioxide in the atmosphere, which also places high demands on reliability. Multiphase flow in porous media is essential to understand CO 2 migration and pressure fields in the subsurface associated with GCS. However, numerical simulation for such problems in 4D is computationally challenging and expensive, due to the multiphysics and multiscale nature of the highly nonlinear governing partial differential equations (PDEs). It prevents us from considering multiple subsurface scenarios and conducting real-time optimization. Here, we develop a Fourier-enhanced multiple-input neural operator (Fourier-MIONet) to learn the solution operator of the problem of multiphase flow in porous media. Fourier-MIONet utilizes the recently developed framework of the multiple-input deep neural operators (MIONet) and incorporates the Fourier neural operator (FNO) in the network architecture. Once Fourier-MIONet is trained, it can predict the evolution of saturation and pressure of the multiphase flow under various reservoir conditions, such as permeability and porosity heterogeneity, anisotropy, injection configurations, and multiphase flow properties. Compared to the enhanced FNO (U-FNO), the proposed Fourier-MIONet has 90% fewer unknown parameters, and it can be trained in significantly less time (about 3.5 times faster) with much lower CPU memory (<15%) and GPU memory (<35%) requirements, to achieve similar prediction accuracy. In addition to the lower computational cost, Fourier-MIONet can be trained with only 6 snapshots of time to predict the PDE solutions for 30 years. Furthermore, we observed that Fourier-MIONet can maintain good accuracy when predicting out-of-distribution (OOD) data. The excellent generalizability of Fourier-MIONet is enabled by its adherence to the physical principle that the solution to a PDE is continuous over time. Furthermore, the developed Fourier-MIONet makes it possible to solve the long-time evolution of geological carbon sequestration in a large-scale three-dimensional space accurately and efficiently.

97 MATHEMATICS AND COMPUTING↗

Extreme-scale EV charging infrastructure planning for last-mile delivery using high-performance parallel computing

Here, this paper addresses stochastic charger location and allocation problems under queue congestion for last-mile delivery using electric vehicles (EVs). The objective is to decide where to open charging stations and how many chargers of each type to install, subject to budgetary and waiting-time constraints. We formulate the problem as a mixed-integer non-linear program, where each station-charger pair is modeled as a multiserver queue with stochastic arrivals and service times to capture the notion of waiting in fleet operations. The model is extremely large, with billions of variables and constraints for a typical metropolitan area; even loading the model in solver memory is difficult, let alone solving it. To address this challenge, we develop a Lagrangian-based dual decomposition framework that decomposes the problem by station and leverages parallelization on high-performance computing systems, where the subproblems are solved by using a cutting plane method and their solutions are collected at the master level. We also develop a three-step rounding heuristic to transform the fractional subproblem solutions into feasible integral solutions. Computational experiments on data from the Chicago metropolitan area with hundreds of thousands of households and thousands of candidate stations show that our approach produces high-quality solutions in cases where existing exact methods cannot even load the model in memory. We also analyze various policy scenarios, demonstrating that combining existing depots with newly built stations under multiagency collaboration substantially reduces costs and congestion. These findings offer a scalable and efficient framework for developing sustainable large-scale EV charging networks.

Capacity allocation↗

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits↗

Quantum Imaging of Ferromagnetic van der Waals Magnetic Domain Structures at Ambient Conditions

Recently discovered 2D van der Waals magnetic materials, and specifically iron–germanium–telluride (Fe5GeTe2), have attracted significant attention both from a fundamental perspective and for potential applications. Key open questions concern their domain structure and magnetic phase transition temperature as a function of sample thickness and external field, as well as implications for integration into devices such as magnetic memories and logic. Here we address key questions using a nitrogen-vacancy center based quantum magnetic microscope, enabling direct imaging of the magnetization of Fe5GeTe2 at submicrometer spatial resolution as a function of temperature, magnetic field, and thickness. This quantum imaging technique provides noninvasive, high-sensitivity measurements with high spatial resolution under ambient conditions, making it particularly well suited for probing 2D magnets. We employ spatially resolved measures, including magnetization variance and cross-correlation, and find a significant spread in transition temperature yet with no clear dependence on thickness down to 15 nm. We also identify previously unknown stripe features in the optical as well as magnetic images, which we attribute to modulations of the constituting elements during crystal synthesis and subsequent oxidation. Our results suggest that the magnetic anisotropy in this material does not play a crucial role in their magnetic properties, leading to a magnetic phase transition of Fe5GeTe2 which is largely thickness-independent down to 15 nm. Our findings could be significant in designing future spintronic devices, magnetic memories, and logic with 2D van der Waals magnetic materials.

Bindu, Bindu [Hebrew University of Jerusalem, Isra↗

Tunable Interfacial to Filamentary Resistive Switching Mechanism in Room-Temperature-Grown Amorphous YBa 2 Cu 3 O x with Excess Cu Addition

Resistive switching technologies have the potential not only to create large efficiency gains in computer memory but also to revolutionize emerging fields such as neuromorphic computing. In this paper, we report on novel resistive switching behavior in devices made from room-temperature-grown Cu-rich amorphous YBa 2 Cu 3 O x (YBCO) films, a material otherwise well-known as a high-temperature superconductor. In Nb:STO substrate/amorphous YBCO film (≈200 nm)/metallic Cu (15 nm)/metallic Pt (15 nm) devices, we demonstrate that the resistive switching can be tuned between mechanisms involving extended areas of the YBCO/electrode interface and a single-point filamentary mechanism simply by changing the Cu content of the deposition target and hence in the films. Changing the Cu content can also be used to optimize the properties of the devices further, with devices with an added 15 mol % of Cu in YBCO initially providing an on/off ratio >100, switching endurance potential >6500 cycles, and state retention >2 × 10 4 s, all at low switching fields of 0.3 MV/cm. The amalgam of promising resistive switching properties, fast growth (150 nm/min) at room temperature, and tuneability of the switching mechanism indicates the strong potential of this proof-of-concept amorphous system for future memory applications.

Cu↗

Autonomous Multistate Nanoencoding Using Combinatorial Ferroelectric Closure Domains in BiFeO 3

Recent advances in ferroic materials have identified topological defects as promising candidates for enabling additional functionalities in future electronic systems. The generation of stable and customizable polar topologies is needed to achieve multistates that enable beyond-binary device architectures. Here, in this study, we show how to autonomously pattern on-demand highly tunable striped closure domains in pristine rhombohedral-phase BiFeO 3 thin films through precise scanning of a biased atomic force microscopy tip along carefully designed paths. By employing this strategy, we generate and manipulate closed-loop structures with high spatial resolution in an automated manner, allowing the creation of highly tunable and intricate topological domain structures that exhibit distinct polarization configurations without the need for electrode deposition or complex heterostructure growth. As a proof-of-concept for ferroelectric beyond-binary memory devices, we use such topological domains as multistates, engineering an alphabet and automating the symbolic writing/reading process using autonomous microscopy. The resulting information density is compared with that of current commercially available memory devices, demonstrating the potential of ferroelectric topological domains for multistate information storage applications.

BiFeO3↗

Anodized Aluminum Oxide Membrane Ionic Memristors

Memory effect in ion transport (IT) at the solid–solution interface is uniquely attractive in that the conductance depends on or “memorizes” the previous states. Hysteretic and rectified transport properties offer exciting potential to developing advanced iontronics and neuromorphic functions, improving the efficiency of energy conversion and electrochemical processes, and overcoming the selectivity-throughput bottleneck in the enrichment of low abundant species for environment- and energy-friendly separations, among others. Herein, memory effects are discovered in the rectified electrokinetic IT through anodized aluminum oxide (AAO) membranes containing densely packed highly ordered nanochannels (10 10 per cm 2 ). Characteristic memristor responses of pinched current–potential loops are resolved in voltammetric experiments and successfully reproduced through finite element simulation. Excitatory and inhibitory conductance states are shown to arise from the enrichment and depletion of mobile charge carriers. Structurewise, the transport symmetry is broken by the barrier oxide layer (BOL) on the one end of the cylindrical nanochannels across the AAO membranes. Charge selectivity is attributed to the gradient(s) of the space charge density across the BOL characterized by depth profiling via X-ray photoelectron spectroscopy analysis. The space charge gradient(s) overcomes the fundamental limitation of widely exploited surface charge effects to enable intense rectification and hysteresis prevailing at very high ionic concentrations up to 1–2 M. A new strategy is developed for controlling the preferential IT direction and selectivity via counterion intercalation and extraction/exchange. Mechanistic understanding is further confirmed through parameter variations such as potential scan rate and ionic strength, which also demonstrates convenient controls of the related functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Mass Conservation Relaxed (MCR) LSTM Model for Streamflow Simulation Across CONUS

The recent development of the physics-aware Mass-Conserving Long Short-Term Memory network (MC-LSTM) provides an alternative to other data-driven Deep Learning (DL) models in hydrology. Mass-Conserving Long Short-Term Memory incorporates mass conservation directly into the LSTM architecture. Despite the theoretical advancements, studies have reported a surprisingly limited performance of the MC-LSTM in streamflow simulation. We hypothesize that such a limitation is due to the unrealistic mass conservation scheme in MC-LSTM, which overlooks unobserved incoming water fluxes beyond precipitation. As an attempt to verify this hypothesis, we propose a Mass Conservation Relaxed LSTM (MCR-LSTM), which incorporates a bi-directional mass relaxation (MR) component to account for potential incoming water fluxes beyond precipitation. We train and test the proposed MCR-LSTM model across 531 watersheds in the contiguous United States (CONUS) against three baseline models: the Sacramento Soil Moisture Accounting, LSTM, and MC-LSTM. Our results show that MCR-LSTM outperforms MC-LSTM despite its underperformance compared to LSTM. Specifically, MCR-LSTM's advantage over MC-LSTM is mainly seen in the Plains and Western U.S., where the newly incorporated MR component better simulates water loss and suggests the likely existence of additional incoming water fluxes beyond precipitation, respectively. The novelty and contribution of this study are twofold: firstly, it introduces an alternative physics-aware DL tool (i.e., MCR-LSTM) in hydrology with higher accuracy in specific regions compared to MC-LSTM. Secondly, it provides a diagnosis of regions where strict, precipitation-based mass conservation constraints may be unrealistic in streamflow simulation.

deep learning↗

Numerically exact configuration interaction at quadrillion-determinant scale

The combinatorial growth of configuration interaction (CI) has long limited this formally exact quantum chemistry method to only the smallest molecules. Here, we report a numerically exact CI calculation exceeding one quadrillion (10 15 ) determinants, made possible by a lossless categorical compression strategy within the small-tensor-product distributed active space (STP-DAS) framework. This approach overcomes the traditional memory bottlenecks of CI by a numerically exact compression of the wavefunction representation and reformulating the most computationally demanding matrix–vector operations. Using this method, we performed a fully relativistic CI calculation of the ground state of HBrTe with over 10 15 complex-valued determinants in just 34.5 h on 1000 computing nodes—the largest CI calculation ever reported. We further achieved fast computation for systems with hundreds of billions of determinants on only a few compute nodes. Extensive benchmarks confirm that the method retains full numerical exactness while cutting memory and computational cost by orders of magnitude. Compared to previous state-of-the-art CI calculations, this work achieves a 1000 times increase in CI space, a 10 6 -fold increase in floating-point operations performed, and a 10 6 -fold improvement in computational speed.

Computational chemistry↗

Spin-optomechanical cavity interfaces by deep subwavelength phonon-photon confinement

A central goal of quantum information science is transferring qubits between space, time, and modality. Spin-based systems in solids are promising quantum memories, but high-fidelity transfer of their quantum states to telecom optical fields remains challenging. Here, we introduce a phonon-mediated interface between spins in a diamond nanobeam optomechanical crystal and telecom optical fields by a simultaneous deep-subwavelength confinement of optical and acoustic fields with mode volumes $V_{\textrm{mech}}$$/Λ^3_\textrm{p} ~ 10^{-5}$ and $V_{\textrm{opt}}$$/λ^3 ~ 10^{−3}$, respectively. This confinement boosts the spin-mechanical coupling rate of Group-IV silicon vacancy (SiV − ) centers by an order of magnitude to ~ 32 MHz while retaining high acousto-optical couplings. The optical cavity couples to the spin irrespective of the emitter’s native excited states, avoiding spectral diffusion. Using Quantum Monte Carlo simulations, we estimate heralded entanglement fidelities exceeding 0.96 between two such interfaces. We anticipate broad utility beyond diamond emitter-telecom systems to most solid-state quantum memories.

Raniwala, Hamza [Massachusetts Inst. of Technology↗

Response of hypoxia to future climate change is sensitive to methodological assumptions

Climate-induced changes in hypoxia are among the most serious threats facing estuaries, which are among the most productive ecosystems on Earth. Future projections of estuarine hypoxia typically involve long-term multi-decadal continuous simulations or more computationally efficient time slice and delta methods that are restricted to short historical and future periods. We make a first comparison of these three methods by applying a linked terrestrial–estuarine model to the Chesapeake Bay, a large coastal-plain estuary in the eastern United States. Results show that the time slice approach accurately captures the behavior of the continuous approach, indicating a minimal impact of model memory. However, increases in mean annual hypoxic volume by the mid-twenty-first century simulated by the delta approach (+ 19%) are approximately twice as large as the time slice and continuous experiments (+ 9% and + 11%, respectively), indicating an important impact of changes in climate variability. Our findings suggest that system memory and projected changes in climate variability, as well as simulation length and natural variability of system hypoxia, should be considered when deciding to apply the more computationally efficient delta and time slice methods.

54 ENVIRONMENTAL SCIENCES↗

Teacher-student training improves the accuracy and efficiency of machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) are revolutionizing the field of molecular dynamics (MD) simulations. Recent MLIPs have tended towards more complex architectures trained on larger datasets. The resulting increase in computational and memory costs may prohibit the application of these MLIPs to perform large-scale MD simulations. Herein, we present a teacher-student training framework in which the latent knowledge from the teacher (atomic energies) is used to augment the students' training. We show that the light-weight student MLIPs have faster MD speeds at a fraction of the memory footprint compared to the teacher models. Remarkably, the student models can even surpass the accuracy of the teachers, even though both are trained on the same quantum chemistry dataset. Our work highlights a practical method for MLIPs to reduce the resources required for large-scale MD simulations.

36 MATERIALS SCIENCE↗

Signal propagation in reversible digital mechanics

Digital mechanics explores information processing through binary, mechanical circuits. This work demonstrates a flexural, mechanical integrated circuit (m-IC) that achieves reversible, non-reciprocal signal propagation through integrated AND logic and memory. Our approach exploits sequential bistable transitions with symmetric energy wells, tunable stiffness, impedance matching, and AND gate non-linearity, to enable signal propagation, repeatability, and reversibility. We present a generalized model of logic kinematics and energetics, validated experimentally, to study energy flows, quantify energetic limits, and identify operating regimes for reversible logic. Macro-scale experiments confirm propagation dynamics, and new fabrication methods extend the architecture to micro-scale devices. By achieving controlled, reversible signal transmission across interconnected logic and memory, this work establishes a scalable platform for robust mechanical computing and adaptive sensing.

Johnson, Hilary A. [Lawrence Livermore National La↗

Isolation of individual Er quantum emitters in anatase TiO 2 on Si photonics

Defects and dopant atoms in solid state materials are a promising platform for realizing single photon sources and quantum memories, which are the basic building blocks of quantum repeaters needed for long distance quantum networks. In particular, trivalent erbium (Er 3+ ) is of interest because it couples C-band telecom optical transitions with a spin-based memory platform. In order to produce quantum repeaters at the scale required for quantum networks it is imperative to integrate these necessary building blocks with mature and scalable semiconductor processes. Here, in this work, we demonstrate the optical isolation of single Er 3+ ions in CMOS-compatible titanium dioxide (TiO 2 ) thin films monolithically integrated on a silicon-on-insulator photonics platform. Our results demonstrate an initial step toward the realization of a monolithically integrated and scalable quantum photonics package based on Er 3+ doped thin films.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field↗

Tree tensor network hierarchical equations of motion based on time-dependent variational principle for efficient open quantum dynamics in structured thermal environments

In this work, we introduce an efficient method, TTN-HEOM, for exactly calculating the open quantum dynamics for driven quantum systems interacting with highly structured bosonic baths by combining the tree tensor network (TTN) decomposition scheme with the bexcitonic generalization of the numerically exact hierarchical equations of motion (HEOM). The method yields a series of quantum master equations for all core tensors in the TTN that efficiently and accurately capture the open quantum dynamics for non-Markovian environments to all orders in the system–bath interaction. These master equations are constructed based on the time-dependent Dirac–Frenkel variational principle, which isolates the optimal dynamics for the core tensors given the TTN ansatz. The dynamics converges to the HEOM when increasing the rank of the core tensors, a limit in which the TTN ansatz becomes exact. We introduce TENSO, tensor equations for non-Markovian structured open systems, as a general-purpose Python code to propagate the TTN-HEOM dynamics. We implement three general propagators for the coupled master equations: two fixed-rank methods that require a constant memory footprint during the dynamics and one adaptive-rank method with a variable memory footprint controlled by the target level of computational error. We exemplify the utility of these methods by simulating a two-level system coupled to a structured bath containing one Drude–Lorentz component and eight Brownian oscillators, which is beyond what can presently be computed using the standard HEOM. Our results show that the TTN-HEOM is capable of simulating both dephasing and relaxation dynamics of driven quantum systems interacting with structured baths, even those of chemical complexity, with an affordable computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On the discretization error of the discrete generalized quantum master equation

The transfer tensor method (TTM) [Cerrillo and Cao, Phys. Rev. Lett. 112 , 110401 (2014)] can be considered a discrete-time formulation of the Nakajima–Zwanzig quantum master equation (NZ-QME) for modeling non-Markovian quantum dynamics. A recent paper [Makri, J. Chem. Theory Comput. 21 , 5037 (2025)] raised concerns regarding the consistency of the TTM discretization, particularly a spurious term at the initial time t = 0. Here, this work presents a detailed analysis of the discretization structure of the TTM, clarifying the origin of the initial-time correction and establishing a consistent relationship between the TTM discrete-time memory kernel K N and the continuous-time NZ-QME kernel $\mathscr{K}$( N Δ t ). This relationship is validated numerically using the spin-boson model, demonstrating convergence of reconstructed memory kernels and accurate dynamical evolution as Δ t → 0. While the TTM provides a consistent discretization, we note that alternative schemes are also viable, such as the midpoint derivative/midpoint integral scheme proposed in Makri’s work. The relative performance of various schemes for either computing accurate $\mathscr{K}$( N Δ t ) from exact dynamics or obtaining accurate dynamics from exact $\mathscr{K}$( N Δ t ) warrants further investigation.

Density-matrix↗