Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

LibraryX: A Framework for Cross-Library-Call Optimization

Scientific applications utilize performance libraries as a software engineering concept: these libraries encapsulate important and well-understood (mathematical) operations, allow for reuse, and are implemented and tuned by experts. Domain scientists then implement complex algorithms based on these domainspecific libraries. While individual library calls are optimized, larger performance gains across sequences of calls—sometimes spanning multiple libraries—are often unrealized, forcing a trade-off between performance and implementation complexity.To overcome this issue, we propose LibraryX, an approach and a system that allows for cross-library-call optimization even when library calls stem from multiple performance libraries. LibraryX annotates library calls with semantic information and optimizes entire directed acyclic graphs (DAGs) of calls dynamically using the SPIRAL code generation system. We demonstrate its effectiveness across a range of memory bound workloads, achieving significant speedups on Nvidia, AMD, and Intel accelerators compared to code using native libraries without cross-call optimization.

Rao, Sanil [Carnegie Mellon University,Department ↗

Structure-preserving neural networks for the regularized entropy-based closure of a linear, kinetic, radiative transport equation

The main challenge of large-scale numerical simulation of radiation transport is the high memory and computation time requirements of discretization methods for kinetic equations. In this work, we derive and investigate a neural network-based approximation to the entropy-based closure method to accurately compute the solution of the multi-dimensional moment system with a low memory footprint and competitive computational time. We extend methods developed for the standard entropy-based closure to the regularized entropy-based closures. The main idea is to interpret structure-preserving neural network approximations of the regularized entropy-based closure as a two-stage approximation to the original entropy-based closure. We conduct a numerical analysis of this approximation and investigate optimal parameter choices. Our numerical experiments demonstrate that the method has a much lower memory footprint than traditional methods with competitive computation times and simulation accuracy. The code and all trained networks are provided on GitHub.

entropy closure↗

Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee

Currently, the Weather Research and Forecasting model (WRF) utilizes shared memory (OpenMP) and distributed memory (MPI) parallelisms. To take advantage of GPU resources on the Perlmutter supercomputer at NERSC, we port parts of the computationally expensive routine Fast Spectral Bin Microphysics (FSBM) to NVIDIA GPUs using OpenMP device offloading directives. To facilitate this process, we explore a workflow for optimization which uses both runtime profilers and a static code inspection tool Codee to refactor the subroutine. We observe an 2.24x overall speedup for the CONUS-12km storm test case.

Wichitrnithed, Chayanon (Namo) [Odin Institute]↗

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (↗

OPAD-EDIFIS Real-Time Processing

The Optical Plume Anomaly Detection (OPAD) detects engine hardware degradation of flight vehicles through identification and quantification of elemental species found in the plume by analyzing the plume emission spectra in a real-time mode. Real-time performance of OPAD relies on extensive software which must report metal amounts in the plume faster than once every 0.5 sec. OPAD software previously written by NASA scientists performed most necessary functions at speeds which were far below what is needed for real-time operation. The research presented in this report improved the execution speed of the software by optimizing the code without changing the algorithms and converting it into a parallelized form which is executed in a shared-memory multiprocessor system. The resulting code was subjected to extensive timing analysis. The report also provides suggestions for further performance improvement by (1) identifying areas of algorithm optimization, (2) recommending commercially available multiprocessor architectures and operating systems to support real-time execution and (3) presenting an initial study of fault-tolerance requirements.

Katsinis, Constantine↗

In-situ TEM EELS analysis of memristive thin films for neuromorphic computing

Neuromorphic computing stands as a promising frontier for advancing AI algorithms and applications like ChatGBT, offering significant energy efficiency gains. This paper delves into the hardware design intricacies of memristive thin films and their elementary switching mechanisms, including anion migration, electron migration, and phase transitions. Through comprehensive analysis of electron energy loss spectroscopy (EELS) data via in-situ transmission electron microscopy (TEM), we will deduce the primary memristive switching mechanisms vital for optimizing thin film fabrication parameters and achieving desired film thickness, conductivity, and memory retention. A single crystal ptype Si substrate was used with TiN as the bottom metal electrode, TiO x as the insulating dielectric layer, and Pt as the top metal electrode. In-situ TEM was able to tell us the thin film didn’t behave like a filamentary or phase transition material. EELS data deduced that electron trapping/detrapping was one of the primary switching mechanisms. By shedding light on these elementary mechanisms, our study aims to catalyze the development of more 2 efficient and effective neuromorphic computing systems to be deployed into mainstream technologies.

97 MATHEMATICS AND COMPUTING↗

RFI channels, 2

The cutoff parameters for a class of channel models exhibiting burst noise behavior were calculated and the performance of interleaved coding strategies was evaluated. It is concluded that, provided the channel memory is large enough and is properly exploited, interleaved coding is nearly optimal.

Mceliece, R. J.↗

Approach range and velocity determination using laser sensors and retroreflector targets

A laser docking sensor study is currently in the third year of development. The design concept is considered to be validated. The concept is based on using standard radar techniques to provide range, velocity, and bearing information. Multiple targets are utilized to provide relative attitude data. The design requirements were to utilize existing space-qualifiable technology and require low system power, weight, and size yet, operate from 0.3 to 150 meters with a range accuracy greater than 3 millimeters and a range rate accuracy greater than 3 mm per second. The field of regard for the system is +/- 20 deg. The transmitter and receiver design features a diode laser, microlens beam steering, and power control as a function of range. The target design consists of five target sets, each having seven 3-inch retroreflectors, arranged around the docking port. The target map is stored in the sensor memory. Phase detection is used for ranging, with the frequency range-optimized. Coarse bearing measurement is provided by the scanning system (one set of binary optics) angle. Fine bearing measurement is provided by a quad detector. A MIL-STD-1750 A/B computer is used for processing. Initial test results indicate a probability of detection greater than 99 percent and a probability of false alarm less than 0.0001. The functional system is currently at the MIT/Lincoln Lab for demonstration.

Donovan, William J.↗

Efficient Calculation of a Jitter/Stability Metric

A tool for computing a jitter/stability metric used in NASA requirements statements is developed. An efficient algorithm is given for computing this metric. Two ways of implementing it on a computer are discussed. One is optimized for computational speed while the other sacrifices some speed to conserve memory. Timing studies are given to show that the improvement of computation times using the present algorithm over previously existing techniques can run to several orders of magnitude, and that previous techniques were so costly that the present algorithm represents enabling technology. Further comparisons show that the memory conservative implementation runs at about half the speed of the fast implementation, but can cut the major data storage requirement of the fast implementation by 95-99%, making the algorithm implementable on much smaller computers, such as PC's, than it would be otherwise. Software for both implementations is included in version 2 of the NASA time and frequency domain analysis program PLATSIM.

Giesy, Daniel P.↗

How efficiently can AI recognize Wireless Devices?

This poster presents a hardware benchmarking methodology for a 3-layer CNN waveform classifier deployed using ONNX Runtime on an NVIDIA Jetson AGX Orin. The dataset consist of 9 signal types, -30 to +30 dB SNR with 5dB increments. Benchmarking on the Jetson AGX Orin gave an accuracy of 91.9% and GPU throughput of 107,120 predictions/sec (23× faster than CPU). The Jetson GPU reached approximately 27M samples/sec with stable performance but fell below the 40 MHz rate needed for real-time radio feeds. Sustained testing of 5 minutes confirmed stable performance with no memory leaks, establishing a reproducible benchmarking baseline for future edge-deployment optimization.

99 - GENERAL AND MISCELLANEOUS↗

Techno-economic implications and cost of forecasting errors in solar PV power production using optimized deep learning models

Accurate solar Photovoltaic (PV) power forecasting is important for enhancing both the performance and economic feasibility of PV systems. This study evaluates several deep learning models, including Dense Neural Networks (DNN), Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), and a hybrid LSTMCNN model, for predicting PV power production one day in advance. Prior to optimization, the models exhibited relatively high errors, with the best model (DNN) achieving a Root Mean Square Error (RMSE) of 31.13 kW and a coefficient of determination (R 2 ) of 62.15 %. After employing Bayesian optimization, the LSTM-CNN model demonstrated the best performance, with the RMSE reduced to 9.79 kW and R 2 improved to 97.62 %, showcasing significant enhancement in predictive accuracy. Here, the economic evaluation considered three cases: rewards for underestimation (0.08 USD/kWh), no rewards, and penalties for both over-and underestimation (120 % of the utility tariff). In the rewards scenario, the LSTM-CNN model reduced the Levelized Cost of Electricity (LCOE) by 4 %, while in the penalty scenario, a backup diesel generator would have increased the LCOE by 49 %. Additionally, the LSTM-CNN model minimized financial losses, achieving the lowest penalties and maximizing net cash flow compared to other models, demonstrating its overall technical and economic superiority.

Deep learning↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Three-dimensional aerodynamic shape optimization using discrete sensitivity analysis

An aerodynamic shape optimization procedure based on discrete sensitivity analysis is extended to treat three-dimensional geometries. The function of sensitivity analysis is to directly couple computational fluid dynamics (CFD) with numerical optimization techniques, which facilitates the construction of efficient direct-design methods. The development of a practical three-dimensional design procedures entails many challenges, such as: (1) the demand for significant efficiency improvements over current design methods; (2) a general and flexible three-dimensional surface representation; and (3) the efficient solution of very large systems of linear algebraic equations. It is demonstrated that each of these challenges is overcome by: (1) employing fully implicit (Newton) methods for the CFD analyses; (2) adopting a Bezier-Bernstein polynomial parameterization of two- and three-dimensional surfaces; and (3) using preconditioned conjugate gradient-like linear system solvers. Whereas each of these extensions independently yields an improvement in computational efficiency, the combined effect of implementing all the extensions simultaneously results in a significant factor of 50 decrease in computational time and a factor of eight reduction in memory over the most efficient design strategies in current use. The new aerodynamic shape optimization procedure is demonstrated in the design of both two- and three-dimensional inviscid aerodynamic problems including a two-dimensional supersonic internal/external nozzle, two-dimensional transonic airfoils (resulting in supercritical shapes), three-dimensional transport wings, and three-dimensional supersonic delta wings. Each design application results in realistic and useful optimized shapes.

Burgreen, Gregory W.↗

Slave finite element for non-linear analysis of engine structures. Volume 2: Programmer's manual and user's manual

The programming aspects of SFENES are described in the User's Manual. The information presented is provided for the installation programmer. It is sufficient to fully describe the general program logic and required peripheral storage. All element generated data is stored externally to reduce required memory allocation. A separate section is devoted to the description of these files thereby permitting the optimization of Input/Output (I/O) time through efficient buffer descriptions. Individual subroutine descriptions are presented along with the complete Fortran source listings. A short description of the major control, computation, and I/O phases is included to aid in obtaining an overall familiarity with the program's components. Finally, a discussion of the suggested overlay structure which allows the program to execute with a reasonable amount of memory allocation is presented.

Witkop, D. L.↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Communication-Aware Orbit Design for Small Spacecraft Swarms around Small Bodies

Exploration of small Solar System bodies has traditionally been performed by single monolithic spacecraft carrying a number of science instruments. However, science instruments typically cannot be operated simultaneously due to the instrument requirements including optimal viewing angle, surface illumination, altitude and ground resolution, power, and data constraints. This observation has motivated interest in multi-spacecraft architectures where a swarm of small spacecraft, each carrying a single science instrument, studies a small body after being deployed by a carrier spacecraft, which then collects data from the vehicles and relays it to Earth. Such architectures hold promise to yield significant improvements in mission efficiency, increases in data quality, and shorter mission duration. A key difficulty in the design of such missions is the selection of orbits for the small spacecraft, which must satisfy not only instrument requirements, but also strict inter-spacecraft communication and on-board storage constraints. To address this, in this paper, we present a novel computationally-efficient optimization algorithm for \emph{communication-aware design} of the orbits of a small spacecraft swarm orbiting a small body. The proposed approach captures constraints including instrument requirements, inter-spacecraft communication bandwidths, and on-board memory usage, and it can accommodate highly irregular gravity field models and surface geometries. We propose an efficient algorithm for optimization of instrument observations and inter-spacecraft communications; we then leverage the differentiable nature of the proposed algorithm to accelerate a gradient-based global search algorithm. Numerical simulations of a six-spacecraft swarm studying 433 Eros show that the proposed approach successfully identifies high-quality orbits, and significantly outperform communication-agnostic optimization techniques, resulting in a 10% increase in scientific returns and a 30% increase in the quality of the collected data.

Rahmani, Amir↗

A high-density magneto-optic memory.

Magneto-optic memory element based on properties of ferrimagnetic garnet with compensation temperature, discussing reading optimization and laser beams intensity

Goldberg, N.↗

Numerical study of sound propagation in a jet flow

An improved computer oriented solution method for problems involving the propagation of sound through a nonuniform jet flow is developed. The method seeks to optimize the use of computer resources such as core storage space and central memory time. Complete formulation details are presented for a jet flow model consisting of a fixed point source on the jet center line in the potential core.

Padula, S. L.↗