Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Surrogate models for development of unconventional shale reservoirs by an integrated numerical approach of hydraulic fracturing, flow and geomechanics, and machine learning

We develop well-completion surrogate models by taking an integrated workflow of hydraulic fracturing, flow, geomechanics, and machine learning simulation. There are three steps in the proposed workflow. First, history-matching processes are conducted with the field data including pumping and production data for characterization. Second, full-physics simulation is performed with various parameters of the field development (e.g., cluster spacing, clusters per stage, pumping rates and times, amount of proppant, and well spacing) to generate multiple simulation results by changing the parameters of the completion design with well-known hydraulic fracturing, reservoir, geomechanics simulators to calculate fracture geometry, reservoir depressurization, induced stress changes. The workflow is demonstrated over a field in the Southern Midland Basin. Here, we take two completion scenarios: a single well case followed by a multi-well case. Finally, a Long Short-Term Memory (LSTM) machine learning algorithm is employed to create surrogate models that can replicate the full-physics simulation results. Furthermore, results show that the trained models applied in the single well and multi-well cases for a particular geological system can provide good accuracy close to those provided by full-physics simulations. Specifically, the site-specific surrogate models can predict fracture parameters (length, height, and surface area) and cumulative production accurately with computational efficiency, suggesting our proposed workflow can be used as a pragmatic tool for expediting the well completion optimization process.

Geomechanics↗

Geometry-aware training of factorized layers in tensor Tucker format

Reducing parameter redundancies in neural network architectures is crucial for achieving feasible computational and memory requirements during train and inference of large networks. Given its easy implementation and flexibility, one promising approach is layer factorization, which reshapes weight tensors into a matrix format and parameterizes it as the product of two rank-r matrices. However, this family of approaches often requires an initial full-model warm-up phase, prior knowledge of a feasible rank, and it is sensitive to parameter initialization.In this work, we introduce a novel approach to train the factors of a Tucker decomposition of the weight tensors. Our training proposal proves to be optimal in locally approximating the original unfactorized dynamics and stable for the initialization. Furthermore, the rank of each mode is dynamically updated during training.We provide a theoretical analysis of the algorithm, showing convergence, approximation and local descent guarantees. The method's performance is further illustrated through a variety of experiments, showing remarkable training compression rates and comparable or even better performance than the full baseline and alternative layer factorization strategies.

Zangrando, Emanuele [Gran Sasso Science Institute ↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

A cell-centered AMR-ALE framework for 3D multi-material hydrodynamics. Part I: Lagrangian and indirect Euler AMR algorithms

Many applications of physics and engineering involve wide ranges of time and spatial scales. The numerical simulation of localized small scales such as shock waves and material interfaces requires a large number of computational cells in these regions. For these applications, Lagrangian and Arbitrary-Lagrangian-Eulerian (ALE) related methods are engaging since the moving mesh feature naturally brings mesh cells on shock discontinuities and material interfaces are carefully captured. In addition, Adaptive-Mesh-Refinement (AMR) strategies aim to optimize computational resources by concentrating finer mesh cells only in areas of interest while using coarser cells elsewhere. A key but challenging AMR requirement consists in efficiently distributing the computational effort to achieve high accuracy without the prohibitive computational costs associated with uniformly fine grids. Here, in this document, the coupling of the p4est AMR library with a cell-centered Lagrangian scheme is presented with the goal to perform reliable 3D Lagrangian-AMR and indirect Euler-AMR multi-material simulations. In particular, it is shown that starting from a 3D indirect ALE code, the memory management and load balancing requirements can be delegated to an external library (here the p4est library) to unlock ALE-AMR capabilities. First, we present a strategy to transcribe the octant-based connectivity of the 3D AMR framework with that of an unstructured mesh of polygonal cells used in Lagrangian hydrodynamics. Then, we show how refinement and coarsening operations must be adapted to the particular Lagrangian framework to ensure the conservation of volume during those steps. Finally, several numerical test cases are presented that demonstrate the capabilities of the Lagrangian-AMR and indirect Euler-AMR algorithms.

3D cell-centered Lagrangian numerical scheme↗

Design and fabrication of robust hybrid photonic crystal cavities

Abstract Heterogeneously integrated hybrid photonic crystal cavities enable strong light–matter interactions with solid state, optically addressable quantum memories. A key challenge to realizing high quality factor ( Q ) hybrid photonic crystals is the reduced index contrast on the substrate compared to suspended devices in air. This challenge is particularly acute for color centers in diamond because of diamond’s high refractive index, which leads to increased scattering loss into the substrate. Here, we develop a design methodology for hybrid photonic crystals utilizing a detailed understanding of substrate-mediated loss, which incorporates sensitivity to fabrication errors as a critical parameter. Using this methodology, we design robust, high-Q, GaAs-on-diamond photonic crystal cavities, and by optimizing our fabrication procedure, we experimentally realize cavities with Q approaching 30,000 at a resonance wavelength of 955 nm.

Abulnaga, Alex↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection

Communication overhead in federated learning (FL) poses a significant challenge for network anomaly detection systems, where the myriad of client configurations and network conditions can severely impact system efficiency and detection accuracy. While existing approaches attempt to address this through individual optimization techniques, they often fail to maintain the delicate balance between reduced overhead and detection performance. This paper presents an adaptive FL framework that dynamically combines batch size optimization, client selection, and asynchronous updates to achieve efficient anomaly detection. Through extensive profiling and experimental analysis on two distinct datasets-UNSW-NBIS for general network traffic and ROAD for automotive networks-our framework reduces communication overhead by 97.6%; (from 700.0s to 16.8s) compared to synchronous baseline approaches while maintaining comparable detection accuracy (95.10%; vs. 95.12%;). Statistical validation using Mann-Whitney U test confirms significant improvements (p < 0.05) over existing FL approaches across both datasets, demonstrating the framework's adaptability to different network security contexts. Detailed profiling analysis reveals the efficiency gains through dramatic reductions in GPU operations and memory transfers while maintaining robust detection performance under varying client conditions.

Marfo, William [University of Texas at El Paso]↗

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗

Unconventional Quantum Advantages for Computation (U-QuAC)

While quantum computing offers the promise of exponential advantages, limited quantum speedups are known, especially for practical applications. To open new avenues for quantum advantages, we propose Unconventional Quantum Advantages for Computation (U-QuACs), with respect to unconventional resources such as space (number of bits or quantum bits of memory required to solve a problem), accuracy of solution, communication, or energy consumption. We focus on space-efficient quantum algorithms, where we seek to design algorithms that solve a problem using much less space than the total size of the input. A natural setting in which space is critical is the streaming model of computation, where the input data arrives sequentially in pieces that must each be processed individually. Streaming is motivated by a variety of problems including analysis of internet traffic or social networks. We design the first exponential quantum space advantage for a natural streaming problem, which also constitutes the first quantum advantage for approximating a discrete optimization problem, albeit with respect to space.

97 MATHEMATICS AND COMPUTING↗

Microstructure-sensitive mechanical behavior of an additively manufactured psuedoelastic shape memory alloy

The additive manufacturing of shape memory alloys into complex geometries enables fabrication of advanced functional systems across a variety of fields and domains. This work presents results focused on the mechanical behavior of additively manufactured shape memory pseudoelastic NiTi. The deformation induced solid state phase transformation from austenite to martensite allows this system to accommodate large recoverable strains. This deformation behavior is fundamentally driven by crystal-scale transformation physics. Laser powder bed fusion processing reveals that the resulting microstructure, both grain morphology and crystallographic texture, is strongly dependent on the manufacturing processing history. Exhaustive mechanical testing demonstrates that these microstructural factors strongly impact both tensile and cyclic stress–strain behavior. Cyclic dissipative behavior, however, is similar across all tested microstructures following an initial transient period. Remarkably, analysis of spatial strain fields during tensile loading reveals two distinctly different localization “modes”. The first is initiation of localized deformation bands which continuously propagate through the tensile bar during loading. In the second mode localization is observed but lacks propagation; instead additional localization cites nucleate during subsequent loading. The latter phenomena is suspected to be driven by grain-scale deformation physics as the localized band morphologies coincide with grain morphologies. These phenomena strongly impact the resulting aggregate stress–strain behavior. Hence, manufacturers and designers of psuedoelastic functional components must at the very least consider the potential variability in properties when considering additive manufacturing processing. More ideally the process–structure–property relations can be used to further tailor and optimize final functional performance.

Additive manufacturing↗

Characterization of throughput on the AXI DMA bus for burst data transfer over Ethernet

cThe Xilinx AXI Direct Memory Access (AXI DMA) module is an efficient solution for medium-speed data transfer in Xilinx SoC FPGAs, supporting data rates greater than 1000 Gbps even in very suboptimal operating modes. It facilitates direct transfer of AXI stream data into processor memory without constant software intervention, which reduces overhead and ensures consistent data logging. By utilizing the FPGA's available memory, large circular buffers (1-5 GiB) are used to buffer data and accommodate network limitations, enabling high-rate data bursts. In this study, we measured the performance of AXI DMA under conditions simulating its lowest practical data transfer speeds. The Arbitrary Length Data Sender was used to transmit AXI stream packets at 32-bit width and 100 MHz frequency, a narrow width and slow speed. Results show that the AXI DMA can transfer up to 3192.76 Mbps with large packet sizes but experiences reduced performance for smaller packets, as low as 2.6 Mbps for 4-byte packets. For Ethernet-limited applications, packet sizes between 8,000 and 16,000 bytes provided optimal transfer speeds of 874 to 1600 Mbps. These findings suggest that the AXI DMA is not the limiting factor in systems where packet sizes exceed 8,000 bytes.

43 PARTICLE ACCELERATORS↗

Object Proxy Patterns for Accelerating Distributed Applications

Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.

Distributed Computing↗

Fourier-MIONet: Fourier-enhanced multiple-input neural operators for multiphase modeling of geological carbon sequestration

Geologic carbon sequestration (GCS) is a safety-critical technology that aims to reduce the amount of carbon dioxide in the atmosphere, which also places high demands on reliability. Multiphase flow in porous media is essential to understand CO 2 migration and pressure fields in the subsurface associated with GCS. However, numerical simulation for such problems in 4D is computationally challenging and expensive, due to the multiphysics and multiscale nature of the highly nonlinear governing partial differential equations (PDEs). It prevents us from considering multiple subsurface scenarios and conducting real-time optimization. Here, we develop a Fourier-enhanced multiple-input neural operator (Fourier-MIONet) to learn the solution operator of the problem of multiphase flow in porous media. Fourier-MIONet utilizes the recently developed framework of the multiple-input deep neural operators (MIONet) and incorporates the Fourier neural operator (FNO) in the network architecture. Once Fourier-MIONet is trained, it can predict the evolution of saturation and pressure of the multiphase flow under various reservoir conditions, such as permeability and porosity heterogeneity, anisotropy, injection configurations, and multiphase flow properties. Compared to the enhanced FNO (U-FNO), the proposed Fourier-MIONet has 90% fewer unknown parameters, and it can be trained in significantly less time (about 3.5 times faster) with much lower CPU memory (<15%) and GPU memory (<35%) requirements, to achieve similar prediction accuracy. In addition to the lower computational cost, Fourier-MIONet can be trained with only 6 snapshots of time to predict the PDE solutions for 30 years. Furthermore, we observed that Fourier-MIONet can maintain good accuracy when predicting out-of-distribution (OOD) data. The excellent generalizability of Fourier-MIONet is enabled by its adherence to the physical principle that the solution to a PDE is continuous over time. Furthermore, the developed Fourier-MIONet makes it possible to solve the long-time evolution of geological carbon sequestration in a large-scale three-dimensional space accurately and efficiently.

97 MATHEMATICS AND COMPUTING↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

A deep learning-based Bayesian framework for high-resolution calibration of building energy models

Calibrating building energy models (BEMs), i.e., closing discrepancy between modeling and field measurements, is of significance to support its applications in building sustainability and resilience analysis. However, as being widely used in practice, current Bayesian calibration is mostly performed in low-resolution (annual or monthly), instead of high-resolution (hourly or sub-hourly), which is crucial to support emerging BEM applications, such as building-renewable energy integration (demand response) and smart control. This is attributable to the gaps in current Bayesian calibration process, including (1) difficulty in supporting reliable high-resolution calibration with over-parameterization and multi-solution issues, (2) inadequacy of meta-model to capture temporal building dynamics in high-resolution, and (3) excessive computational burdens of covariance matrix calculation in Bayesian inference. Therefore, to close these gaps, this research proposes a novel deep learning-based Bayesian calibration framework, involving pre-calibration mechanism, Long Short-Term Memory as surrogate models, and simplified covariance matrix calculation, to calibrate BEMs in high temporal resolution (i.e., hourly) with enhanced accuracy and computational efficiency. Finally, the case study demonstrates its effectiveness to match modeling outcomes with measurements and realize CV-RMSE of < 30 % and NMBE of < 6 % in hourly resolution, as well as a significant reduction of calibration time (by > 99 %, from > 600 h to ~ 1.5 h).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Emulation of Synaptic Plasticity in WO 3 ‐Based Ion‐Gated Transistors

Neuromorphic systems, inspired by the human brain, promise significant advancements in computational efficiency and power consumption by integrating processing and memory functions, thereby addressing the von Neumann bottleneck. This paper explores the synaptic plasticity of a WO3-based ion-gated transistor (IGT) in [EMIM][TFSI] and a 0.1 mol L −1 LiTFSI in [EMIM][TFSI] for neuromorphic computing applications. Cyclic voltammetry (CV), transistor characteristics, and atomic force microscopy (AFM) force–distance (FD) profiling analyses reveal that Li + brings about ion intercalation, together with higher mobility and conductance, and slower response time (τ). WO 3 IGTs exhibit spike amplitude-dependent plasticity (SADP), spike number-dependent plasticity (SNDP), spike duration-dependent plasticity (SDDP), frequency-dependent plasticity (FDP), and paired-pulse facilitation (PPF), which are all crucial for mimicking biological synaptic functions and understanding how to achieve different types of plasticity in the same IGT. The findings underscore the importance of selecting the appropriate ionic medium to optimize the performance of synaptic transistors, enabling the development of neuromorphic systems capable of adaptive learning and real-time processing, which are essential for applications in artificial intelligence (AI).

36 MATERIALS SCIENCE↗

Diagnostics of Mixed-State Topological Order and Breakdown of Quantum Memory

Topological quantum memory can protect information against local errors up to finite error thresholds. Such thresholds are usually determined based on the success of decoding algorithms rather than the intrinsic properties of the mixed states describing corrupted memories. Here we provide an intrinsic characterization of the breakdown of topological quantum memory, which both gives a bound on the performance of decoding algorithms and provides examples of topologically distinct mixed states. We employ three information-theoretical quantities that can be regarded as generalizations of the diagnostics of ground-state topological order, and serve as a definition for topological order in error-corrupted mixed states. We consider the topological contribution to entanglement negativity and two other metrics based on quantum relative entropy and coherent information. In the concrete example of the two-dimensional (2D) Toric code with local bit-flip and phase errors, we map three quantities to observables in 2D classical spin models and analytically show they all undergo a transition at the same error threshold. This threshold is an upper bound on that achieved in any decoding algorithm and is indeed saturated by that in the optimal decoding algorithm for the Toric code. Published by the American Physical Society 2024

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Processability and Material Behavior of NiTi Shape Memory Alloys Using Wire Laser-Directed Energy Deposition (WL-DED)

Utilizing additive manufacturing (AM) techniques with shape memory alloys (SMAs) like NiTi shows great promise for fabricating highly flexible and functionally superior 3D metallic structures. Compared to methods relying on powder feedstocks, wire-based additive manufacturing processes provide a viable alternative, addressing challenges such as chemical composition instability, material availability, higher feedstock costs, and limitations on part size while simplifying process development. This study presented a novel approach by thoroughly assessing the printability of Ni-rich Ni55.94Ti (Wt. %) SMA using the wire laser-directed energy deposition (WL-DED) technique, addressing the existing knowledge gap regarding the laser wire-feed metal additive manufacturing of NiTi alloys. For the first time, the impact of processing parameters—specifically laser power (400–1000 W) and transverse speed (300–900 mm/min)—on single-track fabrication using NiTi wires in the WL-DED process was examined. An optimal range of process parameters was determined to achieve high-quality prints with minimal defects, such as wire dripping, stubbing, and overfilling. Building upon these findings, we printed five distinct cubes, demonstrating the feasibility of producing nearly porosity-free specimens. Notably, this study investigated the effect of energy density on the printed part density, impurity pick-up, transformation temperature, and hardness of the manufactured NiTi cubes. The results from the cube study demonstrated that varying energy densities (46.66–70 J/mm3) significantly affected the quality of the deposits. Lower to intermediate energy densities achieved high relative densities (>99%) and favorable phase transformation temperatures. In contrast, higher energy densities led to instability in melt pool shape, increased porosity, and discrepancies in phase transformation temperatures. These findings highlighted the critical role of precise parameter control in achieving functional NiTi parts and offer valuable insights for advancing AM techniques in fabricating larger high-quality NiTi components. Additionally, our research highlighted important considerations for civil engineering applications, particularly in the development of seismic dampers for energy dissipation in structures, offering a promising solution for enhancing structural performance and energy management in critical infrastructure.

Dabbaghi, Hediyeh↗