Search NASA⌕ Search

SEARCH · Search NASA

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

Least squares linear lags and limited memory filters

Pure autoregressive (AR) models which are linear in the short term, that is when a variable can be predicted by linear regression on a limited number of past observations, are discussed. When evenly spaced observations are available, a fixed set of AR coefficients can be calculated independent of the data. For filtering purposes, such a lag structure can be implemented recursively with an efficient algorithm. The method of computing variance recursively is also derived. A complete algorithm is presented in the appendix.

Discenza, Joseph H.↗

Developing And Scaling an OpenFOAM Model to Study Turbulent Flow in a HFIR Coolant Channel

Improving the understanding of how computational fluid dynamics (CFD) direct numerical simulations (DNS) of flows in the High Flux Isotope Reactor (HFIR) perform when run in parallel using the high performance computing (HPC) platform Summit at the Oak Ridge Leadership Computing Facility (OLCF) is of particular importance to boost the computational tools used to support HFIR conversion to low enriched fuel (LEU). Evaluation of scaling performance was driven by the increasing importance of graphics processing unit (GPU) usage in HPC, which is becoming the standard for modern supercomputers such as Summit. The desired results are to obtain a strong positive correlation between the computational resources dedicated to a problem and the relative speed-up of the simulation in comparison to a benchmark. This capability will allow substantially improvement in HFIR flow analytical capabilities, specifically when predicting turbulence properties at high Reynolds numbers. The study leverages previous simulation results performed with code PHASTA (finite element) on HPC platforms Cori (NERSC) and Theta (ALCF) [1] with computing options provided in the computing platform OpenFOAM (finite volume) at OLCF. Transitioning from PHASTA to OpenFOAM will (1) eliminate dependence on third-party software for mesh generation and manipulation, (2) reduce resource needs by employing modern architectures, and (3) build expertise for future modeling of HFIR-specific problems like heat transfer in involute geometry, entrance effects, flow structure in channel corners, and so on—all important issues when defining the available thermal margins in the transition to LEU. CPUs and GPUs differ significantly in their architecture and utilization, as discussed in the literature [2]. The most important differences are in the approach to computations and their memory. A single GPU contains a large quantity of cores, enabling it to perform with a much higher throughput than a CPU, but execution requires a different approach. GPU codes execute instructions using the Single-Instruction Multiple-Thread (SIMT) approach in which a single instruction is used for groups of threads called warps. A warp typically consists of 32 threads which must execute the same set of instructions, although on separate threads. Alternately, a CPU has far fewer cores that are much more flexible in their operation, excelling at quickly performing more complex serial computations. This is why GPUs have greater throughput when properly utilized. The second important difference is seen when comparing their memory spaces. Limited memory allocations and CPU–GPU communications cause a significant bottleneck in GPU-accelerated programs. Further study was required to properly take advantage of GPU resources. A comprehensive analysis of code performance and the model-specific features of turbulence constitutes the core of this work. In this study, a DNS simulation of HFIR channel turbulence was performed with the finite volume CFD code OpenFOAM v2112 and CUDA v11.0 on Red Hat Enterprise Linux v8.2. The OpenFOAM installation had AMGx integrated to enable GPU acceleration and utilizes the PETSc4FOAM library. The computational resources and the problem size were scaled on CPU and CPU + GPU architectures to gain a better understanding of the performance of a DNS problem on modern computing hardware. The study aimed to analyze the scaling of the code exclusively on CPUs and then to examine the scaling of the codes with GPU acceleration enabled. Scaling studies included CPU and GPU acceleration on a mesh of varying resolution to analyze the impact of problem size relative to computational resources. In the course of preparing the GPU configuration on Summit, mainly using the AMGX solvers, difficulties were encountered stemming from constant changes resulting from extensive ongoing development activities and the changing environment. This resulted in the inability to complete the GPU portion of the work. The code was compiled and tested, but production runs to assess acceleration were not performed because the used discretional compute time allocation expired as year-end approached. The Summit HPC platform is scheduled for decommissioning in 2024, making it unattractive for future use with Nvidia-based GPUs. Therefore, the work will be moved onto NERSC machines in FY24. An application was prepared and submitted, and sufficient node-hours were awarded to continue the research in the next calendar year. This report summarizes work performed thus far, which mostly focused on CPU OpenFOAM computing.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Integrating Commercial Off-The-Shelf (COTS) graphics and extended memory packages with CLIPS

This paper addresses the question of how to mix CLIPS with graphics and how to overcome PC's memory limitations by using the extended memory available in the computer. By adding graphics and extended memory capabilities, CLIPS can be converted into a complete and powerful system development tool, on the other most economical and popular computer platform. New models of PCs have amazing processing capabilities and graphic resolutions that cannot be ignored and should be used to the fullest of their resources. CLIPS is a powerful expert system development tool, but it cannot be complete without the support of a graphics package needed to create user interfaces and general purpose graphics, or without enough memory to handle large knowledge bases. Now, a well known limitation on the PC's is the usage of real memory which limits CLIPS to use only 640 Kb of real memory, but now that problem can be solved by developing a version of CLIPS that uses extended memory. The user has access of up to 16 MB of memory on 80286 based computers and, practically, all the available memory (4 GB) on computers that use the 80386 processor. So if we give CLIPS a self-configuring graphics package that will automatically detect the graphics hardware and pointing device present in the computer, and we add the availability of the extended memory that exists in the computer (with no special hardware needed), the user will be able to create more powerful systems at a fraction of the cost and on the most popular, portable, and economic platform available such as the PC platform.

Callegari, Andres C.↗

L-BFGS Class Implementation in C++

This report presents a header-only C++ class implementation of the Limited-memory BroydenFletcher-Goldfarb-Shanno (L-BFGS) algorithm. The L-BFGS method is a general purpose quasi-Netwon optimization method that builds an approximation of the descent direction from consecutive iterate and gradient vectors. The limited-memory aspect stems from the replacement of the N × N approximation matrix of the original BFGS method with M vectors of length N. An example usage of the class is included along with the reference source code.

97 MATHEMATICS AND COMPUTING↗

A Topographical Lidar System for Terrain-Relative Navigation

An imaging lidar system is being developed for use in navigation, relative to the local terrain. This technology will potentially be used for future spacecraft landing on the Moon. Systems like this one could also be used on Earth for diverse purposes, including mapping terrain, navigating aircraft with respect to terrain and military applications. The system has been field-tested aboard a helicopter in the Mojave Desert. When this system was designed, digitizers with sufficient sampling rate (2 GHz) were only available with very limited memory. Also, it was desirable to limit the amount of data to be transferred between the digitizer and the mass storage between individual frames. One of the novelty design features of this system was to design the system around the limited amount of memory of the digitizer. The system is required to operate over an altitude (distance) range from a few meters to approximately 1 km, but for each scan across the full field of view, the digitizer memory is only able to hold data for an altitude range no more than 100 m. Data acquisition methods in support of the limited 100 m wide altitude range are described.

Liebe, Carl Christian↗

Efficiently modeling neural networks on massively parallel computers

Neural networks are a very useful tool for analyzing and modeling complex real world systems. Applying neural network simulations to real world problems generally involves large amounts of data and massive amounts of computation. To efficiently handle the computational requirements of large problems, we have implemented at Los Alamos a highly efficient neural network compiler for serial computers, vector computers, vector parallel computers, and fine grain SIMD computers such as the CM-2 connection machine. This paper describes the mapping used by the compiler to implement feed-forward backpropagation neural networks for a SIMD (Single Instruction Multiple Data) architecture parallel computer. Thinking Machines Corporation has benchmarked our code at 1.3 billion interconnects per second (approximately 3 gigaflops) on a 64,000 processor CM-2 connection machine (Singer 1990). This mapping is applicable to other SIMD computers and can be implemented on MIMD computers such as the CM-5 connection machine. Our mapping has virtually no communications overhead with the exception of the communications required for a global summation across the processors (which has a sub-linear runtime growth on the order of O(log(number of processors)). We can efficiently model very large neural networks which have many neurons and interconnects and our mapping can extend to arbitrarily large networks (within memory limitations) by merging the memory space of separate processors with fast adjacent processor interprocessor communications. This paper will consider the simulation of only feed forward neural network although this method is extendable to recurrent networks.

Farber, Robert M.↗

Large-Scale Optimization with Linear Equality Constraints Using Reduced Compact Representation

For optimization problems with linear equality constraints, we prove that the (1,1) block of the inverse KKT matrix remains unchanged when projected onto the nullspace of the constraint matrix. In this work, we develop reduced compact representations of the limited-memory inverse BFGS Hessian to compute search directions efficiently when the constraint Jacobian is sparse. Orthogonal projections are implemented by a sparse QR factorization or a preconditioned LSQR iteration. In numerical experiments two proposed trust-region algorithms improve in computation times, often significantly, compared to previous implementations of related algorithms and compared to IPOPT.

97 MATHEMATICS AND COMPUTING↗

Graph-based Reversible Evaluation and Tangents Library

GRETL is a C++ library for evaluation, re-evaluation and algorithmic differentiation of functional operations on an arbitrary computational graph with limited memory usage. Similar to popular machine learning frameworks in Python, like PyTorch and JAX, it tracks and stores both operations and output data as functions are evaluated. Once this composition of functions is built up, the entire chain of operations can be back propagated to compute sensitivities of the final result with respect to any number of inputs. In contrast to most machine learning applications, memory usage becomes the bottleneck for back propagation in many physics applications, especially for time-dependent PDEs. Dynamic check pointing becomes essential. An important distinguishing feature of GRETL is its ability to limit the maximum memory usage by automatically dynamic checkpointing the data output for each graph operation (see Wang, Moin, Iaccarino, 2009). During backpropagation, parts of the graph that are no longer in memory are automatically re-evaluated from upstream checkpointed states as needed for derivative sensitivity calculations (or more precisely, for vector-Jacobian products). GRETL is particularly beneficial for applications, such as coupled multi-physics, where deriving adjoint-based sensitivities and managing checkpoint memory across modules becomes onerous. Cases which can be readily handled by the GRETL library include: different time-integration algorithms per physics (e.g., coupled predictor-corrector algorithms, IMEX, etc.), sub-cycling, asynchronous integrators, state dependent timestep sizes, iterative solvers and coupling algorithms, controller algorithms, and more.

Tupek, MichaelR [Lawrence Livermore National Labor↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

Hong-Ou-Mandel interference of single-photon-level pulses stored in independent room-temperature quantum memories

Abstract Quantum repeater networks require independent absorptive quantum memories capable of storing and retrieving indistinguishable photons to perform high-repetition entanglement swapping operations. The ability to perform these coherent operations at room temperature is of prime importance for the realization of scalable quantum networks. We perform Hong-Ou-Mandel (HOM) interference between photonic polarization states and single-photon-level pulses stored and retrieved from two sets of independent room-temperature quantum memories. We show that the storage and retrieval of polarization states from quantum memories does not degrade the HOM visibility for few-photon-level polarization states in a dual-rail configuration. For single-photon-level pulses, we measure the HOM visibility with various levels of background in a single polarization, single-rail QM, and investigate its dependence on the signal-to-background ratio. We obtain an HOM visibility of 43%, compared to the 48% no-memory limit of our set-up. These results allow us to estimate a 33% visibility for polarization qubits under the same conditions. These demonstrations lay the groundwork for future applications using large-scale memory-assisted quantum networks.

Physics↗

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗

Cryogenic energy storage: Standalone design, rigorous optimization and techno-economic analysis

Energy storage allows flexible use and management of excess electricity and intermittently available renewable energy. Cryogenic energy storage (CES) is a promising storage alternative with a high technology readiness level and maturity, but the round-trip efficiency is often moderate and the Levelized Cost of Storage (LCOS) remains high. The complex flowsheets with intricate thermodynamics at cryogenic temperatures as well as the presence of multiple loops and refrigeration cycles pose considerable challenges for rigorous model-based design and optimization of CES systems. We present an optimization strategy that couples rigorous process simulation and Bayesian optimization with flowsheet decomposition and identification of hidden coupling constraints to optimally design standalone CES systems. Further refinement is done via a local search using the limited-memory Broyden–Fletcher–Goldfarb–Shanno algorithm. Here our results indicate that it is possible to achieve more than 52% round-trip efficiency and an LCOS of $153/MWh for a standalone 100 MW/400 MWh CES system limited to short-term storage with daily charging–discharging. However, a detailed techno-economic assessment reveals that the LCOS considering total capital investment may exceed $267/MWh when all direct and indirect costs of installation and operation are considered.

25 ENERGY STORAGE↗