Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

A Mission-Adaptive Variable Camber Flap Control System to Optimize High Lift and Cruise Lift-to-Drag Ratios of Future N+3 Transport Aircraft

Boeing and NASA are conducting a joint study program to design a wing flap system that will provide mission-adaptive lift and drag performance for future transport aircraft having light-weight, flexible wings. This Variable Camber Continuous Trailing Edge Flap (VCCTEF) system offers a lighter-weight lift control system having two performance objectives: (1) an efficient high lift capability for take-off and landing, and (2) reduction in cruise drag through control of the twist shape of the flexible wing. This control system during cruise will command varying flap settings along the span of the wing in order to establish an optimum wing twist for the current gross weight and cruise flight condition, and continue to change the wing twist as the aircraft changes gross weight and cruise conditions for each mission segment. Design weight of the flap control system is being minimized through use of light-weight shape memory alloy (SMA) actuation augmented with electric actuators. The VCCTEF program is developing better lift and drag performance of flexible wing transports with the further benefits of lighter-weight actuation and less drag using the variable camber shape of the flap.

Mission-Adaptive Wing↗

Object Proxy Patterns for Accelerating Distributed Applications

Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.

Distributed Computing↗

Computational Modeling and Experimental Characterization of Martensitic Transformations in Nicoal for Self-Sensing Materials

Fundamental changes to aero-vehicle management require the utilization of automated health monitoring of vehicle structural components. A novel method is the use of self-sensing materials, which contain embedded sensory particles (SP). SPs are micron-sized pieces of shape-memory alloy that undergo transformation when the local strain reaches a prescribed threshold. The transformation is a result of a spontaneous rearrangement of the atoms in the crystal lattice under intensified stress near damaged locations, generating acoustic waves of a specific spectrum that can be detected by a suitably placed sensor. The sensitivity of the method depends on the strength of the emitted signal and its propagation through the material. To study the transition behavior of the sensory particle inside a metal matrix under load, a simulation approach based on a coupled atomistic-continuum model is used. The simulation results indicate a strong dependence of the particle's pseudoelastic response on its crystallographic orientation with respect to the loading direction and suggest possible ways of optimizing particle sensitivity. The technology of embedded sensory particles will serve as the key element in an autonomous structural health monitoring system that will constantly monitor for damage initiation in service, which will enable quick detection of unforeseen damage initiation in real-time and during onground inspections.

Wallace, T. A.↗

The Efficiency and the Scalability of an Explicit Operator on an IBM POWER4 System

We present an evaluation of the efficiency and the scalability of an explicit CFD operator on an IBM POWER4 system. The POWER4 architecture exhibits a common trend in HPC architectures: boosting CPU processing power by increasing the number of functional units, while hiding the latency of memory access by increasing the depth of the memory hierarchy. The overall machine performance depends on the ability of the caches-buses-fabric-memory to feed the functional units with the data to be processed. In this study we evaluate the efficiency and scalability of one explicit CFD operator on an IBM POWER4. This operator performs computations at the points of a Cartesian grid and involves a few dozen floating point numbers and on the order of 100 floating point operations per grid point. The computations in all grid points are independent. Specifically, we estimate the efficiency of the RHS operator (SP of NPB) on a single processor as the observed/peak performance ratio. Then we estimate the scalability of the operator on a single chip (2 CPUs), a single MCM (8 CPUs), 16 CPUs, and the whole machine (32 CPUs). Then we perform the same measurements for a chache-optimized version of the RHS operator. For our measurements we use the HPM (Hardware Performance Monitor) counters available on the POWER4. These counters allow us to analyze the obtained performance results.

Frumkin, Michael↗

A Parallel Pipelined Renderer for the Time-Varying Volume Data

This paper presents a strategy for efficiently rendering time-varying volume data sets on a distributed-memory parallel computer. Time-varying volume data take large storage space and visualizing them requires reading large files continuously or periodically throughout the course of the visualization process. Instead of using all the processors to collectively render one volume at a time, a pipelined rendering process is formed by partitioning processors into groups to render multiple volumes concurrently. In this way, the overall rendering time may be greatly reduced because the pipelined rendering tasks are overlapped with the I/O required to load each volume into a group of processors; moreover, parallelization overhead may be reduced as a result of partitioning the processors. We modify an existing parallel volume renderer to exploit various levels of rendering parallelism and to study how the partitioning of processors may lead to optimal rendering performance. Two factors which are important to the overall execution time are re-source utilization efficiency and pipeline startup latency. The optimal partitioning configuration is the one that balances these two factors. Tests on Intel Paragon computers show that in general optimal partitionings do exist for a given rendering task and result in 40-50% saving in overall rendering time.

Chiueh, Tzi-Cker↗

Fourier-MIONet: Fourier-enhanced multiple-input neural operators for multiphase modeling of geological carbon sequestration

Geologic carbon sequestration (GCS) is a safety-critical technology that aims to reduce the amount of carbon dioxide in the atmosphere, which also places high demands on reliability. Multiphase flow in porous media is essential to understand CO 2 migration and pressure fields in the subsurface associated with GCS. However, numerical simulation for such problems in 4D is computationally challenging and expensive, due to the multiphysics and multiscale nature of the highly nonlinear governing partial differential equations (PDEs). It prevents us from considering multiple subsurface scenarios and conducting real-time optimization. Here, we develop a Fourier-enhanced multiple-input neural operator (Fourier-MIONet) to learn the solution operator of the problem of multiphase flow in porous media. Fourier-MIONet utilizes the recently developed framework of the multiple-input deep neural operators (MIONet) and incorporates the Fourier neural operator (FNO) in the network architecture. Once Fourier-MIONet is trained, it can predict the evolution of saturation and pressure of the multiphase flow under various reservoir conditions, such as permeability and porosity heterogeneity, anisotropy, injection configurations, and multiphase flow properties. Compared to the enhanced FNO (U-FNO), the proposed Fourier-MIONet has 90% fewer unknown parameters, and it can be trained in significantly less time (about 3.5 times faster) with much lower CPU memory (<15%) and GPU memory (<35%) requirements, to achieve similar prediction accuracy. In addition to the lower computational cost, Fourier-MIONet can be trained with only 6 snapshots of time to predict the PDE solutions for 30 years. Furthermore, we observed that Fourier-MIONet can maintain good accuracy when predicting out-of-distribution (OOD) data. The excellent generalizability of Fourier-MIONet is enabled by its adherence to the physical principle that the solution to a PDE is continuous over time. Furthermore, the developed Fourier-MIONet makes it possible to solve the long-time evolution of geological carbon sequestration in a large-scale three-dimensional space accurately and efficiently.

97 MATHEMATICS AND COMPUTING↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

A deep learning-based Bayesian framework for high-resolution calibration of building energy models

Calibrating building energy models (BEMs), i.e., closing discrepancy between modeling and field measurements, is of significance to support its applications in building sustainability and resilience analysis. However, as being widely used in practice, current Bayesian calibration is mostly performed in low-resolution (annual or monthly), instead of high-resolution (hourly or sub-hourly), which is crucial to support emerging BEM applications, such as building-renewable energy integration (demand response) and smart control. This is attributable to the gaps in current Bayesian calibration process, including (1) difficulty in supporting reliable high-resolution calibration with over-parameterization and multi-solution issues, (2) inadequacy of meta-model to capture temporal building dynamics in high-resolution, and (3) excessive computational burdens of covariance matrix calculation in Bayesian inference. Therefore, to close these gaps, this research proposes a novel deep learning-based Bayesian calibration framework, involving pre-calibration mechanism, Long Short-Term Memory as surrogate models, and simplified covariance matrix calculation, to calibrate BEMs in high temporal resolution (i.e., hourly) with enhanced accuracy and computational efficiency. Finally, the case study demonstrates its effectiveness to match modeling outcomes with measurements and realize CV-RMSE of < 30 % and NMBE of < 6 % in hourly resolution, as well as a significant reduction of calibration time (by > 99 %, from > 600 h to ~ 1.5 h).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Efficient Parallel Kernel Solvers for Computational Fluid Dynamics Applications

Distributed-memory parallel computers dominate today's parallel computing arena. These machines, such as Intel Paragon, IBM SP2, and Cray Origin2OO, have successfully delivered high performance computing power for solving some of the so-called "grand-challenge" problems. Despite initial success, parallel machines have not been widely accepted in production engineering environments due to the complexity of parallel programming. On a parallel computing system, a task has to be partitioned and distributed appropriately among processors to reduce communication cost and to attain load balance. More importantly, even with careful partitioning and mapping, the performance of an algorithm may still be unsatisfactory, since conventional sequential algorithms may be serial in nature and may not be implemented efficiently on parallel machines. In many cases, new algorithms have to be introduced to increase parallel performance. In order to achieve optimal performance, in addition to partitioning and mapping, a careful performance study should be conducted for a given application to find a good algorithm-machine combination. This process, however, is usually painful and elusive. The goal of this project is to design and develop efficient parallel algorithms for highly accurate Computational Fluid Dynamics (CFD) simulations and other engineering applications. The work plan is 1) developing highly accurate parallel numerical algorithms, 2) conduct preliminary testing to verify the effectiveness and potential of these algorithms, 3) incorporate newly developed algorithms into actual simulation packages. The work plan has well achieved. Two highly accurate, efficient Poisson solvers have been developed and tested based on two different approaches: (1) Adopting a mathematical geometry which has a better capacity to describe the fluid, (2) Using compact scheme to gain high order accuracy in numerical discretization. The previously developed Parallel Diagonal Dominant (PDD) algorithm and Reduced Parallel Diagonal Dominant (RPDD) algorithm have been carefully studied on different parallel platforms for different applications, and a NASA simulation code developed by Man M. Rai and his colleagues has been parallelized and implemented based on data dependency analysis. These achievements are addressed in detail in the paper.

Sun, Xian-He↗

Emulation of Synaptic Plasticity in WO 3 ‐Based Ion‐Gated Transistors

Neuromorphic systems, inspired by the human brain, promise significant advancements in computational efficiency and power consumption by integrating processing and memory functions, thereby addressing the von Neumann bottleneck. This paper explores the synaptic plasticity of a WO3-based ion-gated transistor (IGT) in [EMIM][TFSI] and a 0.1 mol L −1 LiTFSI in [EMIM][TFSI] for neuromorphic computing applications. Cyclic voltammetry (CV), transistor characteristics, and atomic force microscopy (AFM) force–distance (FD) profiling analyses reveal that Li + brings about ion intercalation, together with higher mobility and conductance, and slower response time (τ). WO 3 IGTs exhibit spike amplitude-dependent plasticity (SADP), spike number-dependent plasticity (SNDP), spike duration-dependent plasticity (SDDP), frequency-dependent plasticity (FDP), and paired-pulse facilitation (PPF), which are all crucial for mimicking biological synaptic functions and understanding how to achieve different types of plasticity in the same IGT. The findings underscore the importance of selecting the appropriate ionic medium to optimize the performance of synaptic transistors, enabling the development of neuromorphic systems capable of adaptive learning and real-time processing, which are essential for applications in artificial intelligence (AI).

36 MATERIALS SCIENCE↗

Diagnostics of Mixed-State Topological Order and Breakdown of Quantum Memory

Topological quantum memory can protect information against local errors up to finite error thresholds. Such thresholds are usually determined based on the success of decoding algorithms rather than the intrinsic properties of the mixed states describing corrupted memories. Here we provide an intrinsic characterization of the breakdown of topological quantum memory, which both gives a bound on the performance of decoding algorithms and provides examples of topologically distinct mixed states. We employ three information-theoretical quantities that can be regarded as generalizations of the diagnostics of ground-state topological order, and serve as a definition for topological order in error-corrupted mixed states. We consider the topological contribution to entanglement negativity and two other metrics based on quantum relative entropy and coherent information. In the concrete example of the two-dimensional (2D) Toric code with local bit-flip and phase errors, we map three quantities to observables in 2D classical spin models and analytically show they all undergo a transition at the same error threshold. This threshold is an upper bound on that achieved in any decoding algorithm and is indeed saturated by that in the optimal decoding algorithm for the Toric code. Published by the American Physical Society 2024

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Processability and Material Behavior of NiTi Shape Memory Alloys Using Wire Laser-Directed Energy Deposition (WL-DED)

Utilizing additive manufacturing (AM) techniques with shape memory alloys (SMAs) like NiTi shows great promise for fabricating highly flexible and functionally superior 3D metallic structures. Compared to methods relying on powder feedstocks, wire-based additive manufacturing processes provide a viable alternative, addressing challenges such as chemical composition instability, material availability, higher feedstock costs, and limitations on part size while simplifying process development. This study presented a novel approach by thoroughly assessing the printability of Ni-rich Ni55.94Ti (Wt. %) SMA using the wire laser-directed energy deposition (WL-DED) technique, addressing the existing knowledge gap regarding the laser wire-feed metal additive manufacturing of NiTi alloys. For the first time, the impact of processing parameters—specifically laser power (400–1000 W) and transverse speed (300–900 mm/min)—on single-track fabrication using NiTi wires in the WL-DED process was examined. An optimal range of process parameters was determined to achieve high-quality prints with minimal defects, such as wire dripping, stubbing, and overfilling. Building upon these findings, we printed five distinct cubes, demonstrating the feasibility of producing nearly porosity-free specimens. Notably, this study investigated the effect of energy density on the printed part density, impurity pick-up, transformation temperature, and hardness of the manufactured NiTi cubes. The results from the cube study demonstrated that varying energy densities (46.66–70 J/mm3) significantly affected the quality of the deposits. Lower to intermediate energy densities achieved high relative densities (>99%) and favorable phase transformation temperatures. In contrast, higher energy densities led to instability in melt pool shape, increased porosity, and discrepancies in phase transformation temperatures. These findings highlighted the critical role of precise parameter control in achieving functional NiTi parts and offer valuable insights for advancing AM techniques in fabricating larger high-quality NiTi components. Additionally, our research highlighted important considerations for civil engineering applications, particularly in the development of seismic dampers for energy dissipation in structures, offering a promising solution for enhancing structural performance and energy management in critical infrastructure.

Dabbaghi, Hediyeh↗

Ion-Conducting Organic/Inorganic Polymers

Ion-conducting polymers that are hybrids of organic and inorganic moieties and that are suitable for forming into solid-electrolyte membranes have been invented in an effort to improve upon the polymeric materials that have been used previously for such membranes. Examples of the prior materials include perfluorosulfonic acid-based formulations, polybenzimidazoles, sulfonated polyetherketone, sulfonated naphthalenic polyimides, and polyethylene oxide (PEO)-based formulations. Relative to the prior materials, the polymers of the present invention offer greater dimensional stability, greater ease of formation into mechanically resilient films, and acceptably high ionic conductivities over wider temperature ranges. Devices in which films made of these ion-conducting organic/inorganic polymers could be used include fuel cells, lithium batteries, chemical sensors, electrochemical capacitors, electrochromic windows and display devices, and analog memory devices. The synthesis of a polymer of this type (see Figure 1) starts with a reaction between an epoxide-functionalized alkoxysilane and a diamine. The product of this reaction is polymerized by hydrolysis and condensation of the alkoxysilane group, producing a molecular network that contains both organic and inorganic (silica) links. The silica in the network contributes to the ionic conductivity and to the desired thermal and mechanical properties. Examples of other diamines that have been used in the reaction sequence of Figure 1 are shown in Figure 2. One can use any of these diamines or any combination of them in proportions chosen to impart desired properties to the finished product. Alternatively or in addition, one could similarly vary the functionality of the alkoxysilane to obtain desired properties. The variety of available alkoxysilanes and diamines thus affords flexibility to optimize the organic/inorganic polymer for a given application.

Kinder, James D.↗

Slow Wave Sleep and Long Duration Spaceflight

While ground research has clearly shown that preserving adequate quantities of sleep is essential for optimal health and performance, changes in the progression, order and /or duration of specific stages of sleep is also associated with deleterious outcomes. As seen in Figure 1, in healthy individuals, REM and Non-REM sleep alternate cyclically, with stages of Non-REM sleep structured chronologically. In the early parts of the night, for instance, Non-REM stages 3 and 4 (Slow Wave Sleep, or SWS) last longer while REM sleep spans shorter; as night progresses, the length of SWS is reduced as REM sleep lengthens. This process allows for SWS to establish precedence , with increases in SWS seen when recovering from sleep deprivation. SWS is indeed regarded as the most restorative portion of sleep. During SWS, physiological activities such as hormone secretion, muscle recovery, and immune responses are underway, while neurological processes required for long term learning and memory consolidation, also occur. The structure and duration of specific sleep stages may vary independent of total sleep duration, and changes in the structure and duration have been shown to be associated with deleterious outcomes. Individuals with narcolepsy enter sleep through REM as opposed to stage 1 of NREM. Disrupting slow wave sleep for several consecutive nights without reducing total sleep duration or sleep efficiency is associated with decreased pain threshold, increased discomfort, fatigue, and the inflammatory flare response in skin. Depression has been shown to be associated with a reduction of slow wave sleep and increased REM sleep. Given research that shows deleterious outcomes are associated with changes in sleep structure, it is essential to characterize and mitigate not only total sleep duration, but also changes in sleep stages.

Whitmire, Alexandra↗

Fabricating Composite-Material Structures Containing SMA Ribbons

An improved method of designing and fabricating laminated composite-material (matrix/fiber) structures containing embedded shape-memory-alloy (SMA) actuators has been devised. Structures made by this method have repeatable, predictable properties, and fabrication processes can readily be automated. Such structures, denoted as shape-memory-alloy hybrid composite (SMAHC) structures, have been investigated for their potential to satisfy requirements to control the shapes or thermoelastic responses of themselves or of other structures into which they might be incorporated, or to control noise and vibrations. Much of the prior work on SMAHC structures has involved the use SMA wires embedded within matrices or within sleeves through parent structures. The disadvantages of using SMA wires as the embedded actuators include (1) complexity of fabrication procedures because of the relatively large numbers of actuators usually needed; (2) sensitivity to actuator/ matrix interface flaws because voids can be of significant size, relative to wires; (3) relatively high rates of breakage of actuators during curing of matrix materials because of sensitivity to stress concentrations at mechanical restraints; and (4) difficulty of achieving desirable overall volume fractions of SMA wires when trying to optimize the integration of the wires by placing them in selected layers only.

Turner, Travis L.↗

Development and Characterization of Embedded Sensory Particles Using Multi-Scale 3D Digital Image Correlation

A method for detecting fatigue cracks has been explored at NASA Langley Research Center. Microscopic NiTi shape memory alloy (sensory) particles were embedded in a 7050 aluminum alloy matrix to detect the presence of fatigue cracks. Cracks exhibit an elevated stress field near their tip inducing a martensitic phase transformation in nearby sensory particles. Detectable levels of acoustic energy are emitted upon particle phase transformation such that the existence and location of fatigue cracks can be detected. To test this concept, a fatigue crack was grown in a mode-I single-edge notch fatigue crack growth specimen containing sensory particles. As the crack approached the sensory particles, measurements of particle strain, matrix-particle debonding, and phase transformation behavior of the sensory particles were performed. Full-field deformation measurements were performed using a novel multi-scale optical 3D digital image correlation (DIC) system. This information will be used in a finite element-based study to determine optimal sensory material behavior and density.

Cornell, Stephen R.↗

Neural correspondence to spectrum of environmental uncertainty in multiple-cue probability judgment system with time delay

Despite state-of-the-art technologies like artificial intelligence, human judgment is critically essential in cooperative systems, such as the multi-agent system (MAS), which collect information among agents based on multiple-cue judgment. Human agents can prevent impaired situational awareness of automated agents by confirming situations under environmental uncertainty. System error caused by uncertainty can result in an unreliable system environment, and this environment affects the human agent, resulting in non-optimal decision-making in MAS. Thus, it is necessary to know how human behavior is changed to capture system reliability under uncertainty. Another issue affecting MAS is time delay, which can delay agent information transfer, resulting in low performance and instability. However, it is difficult to find studies on the influence of time delay on human agents. This study is about understanding the human decision-making process under a specific system reliability environment by uncertainty with time delay. We used concepts of expected and unexpected uncertainty to implement reliability of the system usage environment with three types of time delay conditions: no time delay, regular time delay, and irregular time delay conditions. We used electroencephalogram (EEG) for human cognitive neural mechanisms in multiple-cue judgment systems to understand human decision-making. In the reliability of system usage environment, the unreliable system environment significantly creates less memory load by less utilization of system rules for decision-making. In terms of time delay, delayed information delivery does not significantly affect memory load for decision-making.

cognitive process↗

Genetic Algorithm-Guided, Adaptive Model Order Reduction of Flexible Aircrafts

This paper presents a methodology for automated model order reduction (MOR) of flexible aircrafts to construct linear parameter-varying (LPV) reduced order models (ROM) for aeroservoelasticity (ASE) analysis and control synthesis in broad flight parameter space. The novelty includes utilization of genetic algorithms (GAs) to automatically determine the states for reduction while minimizing the trial-and-error process and heuristics requirement to perform MOR; balanced truncation for unstable systems to achieve locally optimal realization of the full model; congruence transformation for "weak" fulfillment of state consistency across the entire flight parameter space; and ROM interpolation based on adaptive grid refinement to generate a globally functional LPV ASE ROM. The methodology is applied to the X-56A MUTT model currently being tested at NASA/AFRC for flutter suppression and gust load alleviation. Our studies indicate that X-56A ROM with less than one-seventh the number of states relative to the original model is able to accurately predict system response among all input-output channels for pitch, roll, and ASE control at various flight conditions. The GA-guided approach exceeds manual and empirical state selection in terms of efficiency and accuracy. The adaptive refinement allows selective addition of the grid points in the parameter space where flight dynamics varies dramatically to enhance interpolation accuracy without over-burdening controller synthesis and onboard memory efforts downstream. The present MOR framework can be used by control engineers for robust ASE controller synthesis and novel vehicle design.

Numerical Analysi↗