Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Problem size, parallel architecture, and optimal speedup

The communication and synchronization overhead inherent in parallel processing can lead to situations where adding processors to the solution method actually increases execution time. Problem type, problem size, and architecture type all affect the optimal number of processors to employ. The numerical solution of an elliptic partial differential equation is examined in order to study the relationship between problem size and architecture. The equation's domain is discretized into n sup 2 grid points which are divided into partitions and mapped onto the individual processor memories. The relationships between grid size, stencil type, partitioning strategy, processor execution time, and communication network type are analytically quantified. In so doing, the optimal number of processors was determined to assign to the solution, and identified (1) the smallest grid size which fully benefits from using all available processors, (2) the leverage on performance given by increasing processor speed or communication network speed, and (3) the suitability of various architectures for large numerical problems.

Nicol, David M.↗

Portable parallel stochastic optimization for the design of aeropropulsion components

This report presents the results of Phase 1 research to develop a methodology for performing large-scale Multi-disciplinary Stochastic Optimization (MSO) for the design of aerospace systems ranging from aeropropulsion components to complete aircraft configurations. The current research recognizes that such design optimization problems are computationally expensive, and require the use of either massively parallel or multiple-processor computers. The methodology also recognizes that many operational and performance parameters are uncertain, and that uncertainty must be considered explicitly to achieve optimum performance and cost. The objective of this Phase 1 research was to initialize the development of an MSO methodology that is portable to a wide variety of hardware platforms, while achieving efficient, large-scale parallelism when multiple processors are available. The first effort in the project was a literature review of available computer hardware, as well as review of portable, parallel programming environments. The first effort was to implement the MSO methodology for a problem using the portable parallel programming language, Parallel Virtual Machine (PVM). The third and final effort was to demonstrate the example on a variety of computers, including a distributed-memory multiprocessor, a distributed-memory network of workstations, and a single-processor workstation. Results indicate the MSO methodology can be well-applied towards large-scale aerospace design problems. Nearly perfect linear speedup was demonstrated for computation of optimization sensitivity coefficients on both a 128-node distributed-memory multiprocessor (the Intel iPSC/860) and a network of workstations (speedups of almost 19 times achieved for 20 workstations). Very high parallel efficiencies (75 percent for 31 processors and 60 percent for 50 processors) were also achieved for computation of aerodynamic influence coefficients on the Intel. Finally, the multi-level parallelization strategy that will be needed for large-scale MSO problems was demonstrated to be highly efficient. The same parallel code instructions were used on both platforms, demonstrating portability. There are many applications for which MSO can be applied, including NASA's High-Speed-Civil Transport, and advanced propulsion systems. The use of MSO will reduce design and development time and testing costs dramatically.

Sues, Robert H.↗

Arctic cognition: a study of cognitive performance in summer and winter at 69 degrees N

Evidence has accumulated over the past 15 years that affect in humans is cyclical. In winter there is a tendency to depression, with remission in summer, and this effect is stronger at higher latitudes. In order to determine whether human cognition is similarly rhythmical, this study investigated the cognitive processes of 100 participants living at 69 degrees N. Participants were tested in summer and winter on a range of cognitive tasks, including verbal memory, attention and simple reaction time tasks. The seasonally counterbalanced design and the very northerly latitude of this study provide optimal conditions for detecting impaired cognitive performance in winter, and the conclusion is negative: of five tasks with seasonal effects, four had disadvantages in summer. Like the menstrual cycle, the circannual cycle appears to influence mood but not cognition.

Cognition↗

A comparison of multiprocessor scheduling methods for iterative data flow architectures

A comparative study is made between the Algorithm to Architecture Mapping Model (ATAMM) and three other related multiprocessing models from the published literature. The primary focus of all four models is the non-preemptive scheduling of large-grain iterative data flow graphs as required in real-time systems, control applications, signal processing, and pipelined computations. Important characteristics of the models such as injection control, dynamic assignment, multiple node instantiations, static optimum unfolding, range-chart guided scheduling, and mathematical optimization are identified. The models from the literature are compared with the ATAMM for performance, scheduling methods, memory requirements, and complexity of scheduling and design procedures.

Storch, Matthew↗

Joint Japan/U.S. Conference on Adaptive Structures, 2nd, Nagoya, Japan, Nov. 12-14, 1991, Collection of Papers

The present conference discusses the development status of adaptive structures in Europe and in Japan, the 'Cosmo-Lab' structures/robotics cooperation concept, active-adhesion concepts for in-orbit structural assembly, adaptively controlled truss structures, object-oriented modeling in structural analysis, the control effectiveness and energy efficiency of an active mass damper, a space truss with experimental tendon control, and piezoelectric actuator-based space trusses. Also discussed is the control of resonant frequencies in adaptive structures through prestressing, active control of vortex-excited vibrations of flexible cylindrical structures, shape adjustment of a flexible space antenna reflector, the SDIO Adaptive Structures Program, optimal trajectories of iterative manipulation for space robots, a docking device as an adaptive structure, shape-memory polymers and their hybrid composites, and fuzzy control methods for structural dynamics.

Matsuzaki, Yuji↗

NASA Electronic Library System (NELS) optimization

This is a compilation of NELS (NASA Electronic Library System) Optimization progress/problem, interim, and final reports for all phases. The NELS database was examined, particularly in the memory, disk contention, and CPU, to discover bottlenecks. Methods to increase the speed of NELS code were investigated. The tasks included restructuring the existing code to interact with others more effectively. An error reporting code to help detect and remove bugs in the NELS was added. Report writing tools were recommended to integrate with the ASV3 system. The Oracle database management system and tools were to be installed on a Sun workstation, intended for demonstration purposes.

Pribyl, William L.↗

HTMT-class Latency Tolerant Parallel Architecture for Petaflops Scale Computation

Computational Aero Sciences and other numeric intensive computation disciplines demand computing throughputs substantially greater than the Teraflops scale systems only now becoming available. The related fields of fluids, structures, thermal, combustion, and dynamic controls are among the interdisciplinary areas that in combination with sufficient resolution and advanced adaptive techniques may force performance requirements towards Petaflops. This will be especially true for compute intensive models such as Navier-Stokes are or when such system models are only part of a larger design optimization computation involving many design points. Yet recent experience with conventional MPP configurations comprising commodity processing and memory components has shown that larger scale frequently results in higher programming difficulty and lower system efficiency. While important advances in system software and algorithms techniques have had some impact on efficiency and programmability for certain classes of problems, in general it is unlikely that software alone will resolve the challenges to higher scalability. As in the past, future generations of high-end computers may require a combination of hardware architecture and system software advances to enable efficient operation at a Petaflops level. The NASA led HTMT project has engaged the talents of a broad interdisciplinary team to develop a new strategy in high-end system architecture to deliver petaflops scale computing in the 2004/5 timeframe. The Hybrid-Technology, MultiThreaded parallel computer architecture incorporates several advanced technologies in combination with an innovative dynamic adaptive scheduling mechanism to provide unprecedented performance and efficiency within practical constraints of cost, complexity, and power consumption. The emerging superconductor Rapid Single Flux Quantum electronics can operate at 100 GHz (the record is 770 GHz) and one percent of the power required by convention semiconductor logic. Wave Division Multiplexing optical communications can approach a peak per fiber bandwidth of 1 Tbps and the new Data Vortex network topology employing this technology can connect tens of thousands of ports providing a bi-section bandwidth on the order of a Petabyte per second with latencies well below 100 nanoseconds, even under heavy loads. Processor-in-Memory (PIM) technology combines logic and memory on the same chip exposing the internal bandwidth of the memory row buffers at low latency. And holographic storage photorefractive storage technologies provide high-density memory with access a thousand times faster than conventional disk technologies. Together these technologies enable a new class of shared memory system architecture with a peak performance in the range of a Petaflops but size and power requirements comparable to today's largest Teraflops scale systems. To achieve high-sustained performance, HTMT combines an advanced multithreading processor architecture with a memory-driven coarse-grained latency management strategy called "percolation", yielding high efficiency while reducing the much of the parallel programming burden. This paper will present the basic system architecture characteristics made possible through this series of advanced technologies and then give a detailed description of the new percolation approach to runtime latency management.

Sterling, Thomas↗

Lessons Learned in the Flight Qualification of the S-NPP and NOAA-20 Solar Array Mechanisms

Deployable solar arrays are the energy source used on almost all Earth orbiting spacecraft and their release and deployment are mission-critical; fully testing them on the ground is a challenging endeavor. The 8 meter long deployable arrays flown on two sequential NASA weather satellites were each comprised of three rigid panels almost 2 meters wide. These large panels were deployed by hinges comprised of stacked constant force springs, eddy current dampers, and were restrained through launch by a set of four releasable hold-downs using shape memory alloy release devices. The ground qualification testing of such unwieldy deployable solar arrays, whose design was optimized for orbital operations, proved to be quite challenging and provides numerous lessons learned. A paperwork review and follow-up inspection after hardware storage determined that there were negative torque margins and missing lubricant, this paper will explain how these unexpected issues were overcome. The paper will also provide details on how the hinge subassemblies, the fully-assembled array, and mechanical ground support equipment were subsequently improved and qualified for a follow-on flight with considerably less difficulty. The solar arrays built by Ball Aerospace Corp. for the Suomi National Polar Partnership (S-NPP) satellite and the Joint Polar Satellite System (JPSS-1) satellite (now NOAA-20) were both successfully deployed on-obit and are performing well.

Helfrich, Daniel↗

Lessons Learned in the Flight Qualification of the S-NPP and NOAA-20 Solar Array Mechanisms

Deployable solar arrays are the energy source used on almost all Earth orbiting spacecraft and their release and deployment are mission-critical; fully testing them on the ground is a challenging endeavor. The 8 meter long deployable arrays flown on two sequential NASA weather satellites were each comprised of three rigid panels almost 2 meters wide. These large panels were deployed by hinges comprised of stacked constant force springs, eddy current dampers, and were restrained through launch by a set of four releasable hold-downs using shape memory alloy release devices. The ground qualification testing of such unwieldy deployable solar arrays, whose design was optimized for orbital operations, proved to be quite challenging and provides numerous lessons learned. A paperwork review and follow-up inspection after hardware storage determined that there were negative torque margins and missing lubricant, this paper will explain how these unexpected issues were overcome. The paper will also provide details on how the hinge subassemblies, the fully-assembled array, and mechanical ground support equipment were subsequently improved and qualified for a follow-on flight with considerably less difficulty. The solar arrays built by Ball Aerospace Corp. for the Suomi National Polar Partnership (SNPP) satellite and the Joint Polar Satellite System (JPSS-1) satellite (now NOAA-20) were both successfully deployed on-obit and are performing well.

Sexton, Adam↗

Optimal evaluation of array expressions on massively parallel machines

We investigate the problem of evaluating FORTRAN 90 style array expressions on massively parallel distributed-memory machines. On such machines, an elementwise operation can be performed in constant time for arrays whose corresponding elements are in the same processor. If the arrays are not aligned in this manner, the cost of aligning them is part of the cost of evaluating the expression. The choice of where to perform the operation then affects this cost. We present algorithms based on dynamic programming to solve this problem efficiently for a wide variety of interconnection schemes, including multidimensional grids and rings, hypercubes, and fat-trees. We also consider expressions containing operations that change the shape of the arrays, and show that our approach extends naturally to handle this case.

Chatterjee, Siddhartha↗

Spectral methods in time for hyperbolic equations

A pseudospectral numerical scheme for solving linear, periodic, hyperbolic problems is described. It has infinite accuracy both in time and in space. The high accuracy in time is achieved without increasing the computational work and memory space which is needed for a regular, one step explicit scheme. The algorithm is shown to be optimal in the sense that among all the explicit algorithms of a certain class it requires the least amount of work to achieve a certain given resolution. The class of algorithms referred to consists of all explicit schemes which may be represented as a polynomial in the spatial operator.

Tal-Ezer, H.↗

Balancing Contention and Synchronization on the Intel Paragon

The Intel Paragon is a mesh-connected distributed memory parallel computer. It uses an oblivious and deterministic message routing algorithm: this permits us to develop highly optimized schedules for frequently needed communication patterns. The complete exchange is one such pattern. Several approaches are available for carrying it out on the mesh. We study an algorithm developed by Scott. This algorithm assumes that a communication link can carry one message at a time and that a node can only transmit one message at a time. It requires global synchronization to enforce a schedule of transmissions. Unfortunately global synchronization has substantial overhead on the Paragon. At the same time the powerful interconnection mechanism of this machine permits 2 or 3 messages to share a communication link with minor overhead. It can also overlap multiple message transmission from the same node to some extent. We develop a generalization of Scott's algorithm that executes complete exchange with a prescribed contention. Schedules that incur greater contention require fewer synchronization steps. This permits us to tradeoff contention against synchronization overhead. We describe the performance of this algorithm and compare it with Scott's original algorithm as well as with a naive algorithm that does not take interconnection structure into account. The Bounded contention algorithm is always better than Scott's algorithm and outperforms the naive algorithm for all but the smallest message sizes. The naive algorithm fails to work on meshes larger than 12 x 12. These results show that due consideration of processor interconnect and machine performance parameters is necessary to obtain peak performance from the Paragon and its successor mesh machines.

Bokhari, Shahid H.↗

Reduced Navier Stokes Relaxation Procedures for Internal Flows

In spite of significant advancement in the field of high speed computing, flow calculations involving complex geometries and/or flow behavior still require large amounts of CPU time and memory. In order to predict such flows without sacrificing grid convergence and accuracy, adaptive gridding techniques that provide optimal resolution are highly desirable. The present work combines multigrid techniques and domain decomposition concepts to provide local, solution adaptive, grid refinement. Several viscous compressible and incompressible, two and three-dimensional, flows with strong inviscid interaction and/or axial flow reversal, are considered with a segmented multigrid domain decomposition (SMGDD) procedure for which uniform meshes result in each domain. A pressure-based form of flux-vector splitting is applied to the Navier-Stokes equations, which are represented by an implicit lowest-order reduced Navier-Stokes (RNS) system and a purely diffusive, higher-order, deferred-corrector. A trapezoidal or box-like form of discretization insures that all mass conservation properties are satisfied at interfacial and outflow boundaries, even for this primitive-variable non-staggered grid computation. The SMGDD technique presented herein has previously been applied for incompressible two dimensional flows. The present work offers improvement in the gridding strategy, by allowing for disjoint subdomains that provide optimal resolution of disparate flow features. It also extends the SMGDD technique to three dimensional compressible flows. Laminar and turbulent flow in a backward facing step channel is considered; although the procedure is applicable to more severe geometries. The standard K-epsilon model is applied for turbulence closure. For Re greater than 400, differences between two-dimensional theory and experiment are resolved through a three dimensional simulation, which confirms the experimentally observed three dimensionality of the recirculation patterns on the upper and lower surfaces.

Rubin, Stanley G.↗

Parallelization of an Object-Oriented Unstructured Aeroacoustics Solver

A computational aeroacoustics code based on the discontinuous Galerkin method is ported to several parallel platforms using MPI. The discontinuous Galerkin method is a compact high-order method that retains its accuracy and robustness on non-smooth unstructured meshes. In its semi-discrete form, the discontinuous Galerkin method can be combined with explicit time marching methods making it well suited to time accurate computations. The compact nature of the discontinuous Galerkin method also makes it well suited for distributed memory parallel platforms. The original serial code was written using an object-oriented approach and was previously optimized for cache-based machines. The port to parallel platforms was achieved simply by treating partition boundaries as a type of boundary condition. Code modifications were minimal because boundary conditions were abstractions in the original program. Scalability results are presented for the SCI Origin, IBM SP2, and clusters of SGI and Sun workstations. Slightly superlinear speedup is achieved on a fixed-size problem on the Origin, due to cache effects.

Baggag, Abdelkader↗

Error Estimation and h-Adaptivity for Optimal Finite Element Analysis

The objective of adaptive meshing and automatic error control in finite element analysis is to eliminate the need for the application engineer from re-meshing and re-running design simulations to verify numerical accuracy. The user should only need to enter the component geometry and a coarse finite element mesh. The software will then autonomously and adaptively refine this mesh where needed, reducing the error in the fields to a user prescribed value. The ideal end result of the simulation is a measurable quantity (e.g. scattered field, input impedance), calculated to a prescribed error, in less time and less machine memory than if the user applied typical uniform mesh refinement by hand. It would also allow for the simulation of larger objects since an optimal mesh is created.

Cwik, Tom↗

Memory-efficient decoding of LDPC codes

We present a low-complexity quantization scheme for the implementation of regular (3,6) LDPC codes. The quantization parameters are optimized to maximize the mutual information between the source and the quantized messages. Using this non-uniform quantized belief propagation algorithm, we have simulated that an optimized 3-bit quantizer operates with 0.2dB implementation loss relative to a floating point decoder, and an optimized 4-bit quantizer operates less than 0.1dB quantization loss.

Low-Density-Parity-Check (LDPC)↗

Adaptation and optimization of a line-by-line radiative transfer program for the STAR-100 (STARSMART)

A program to calculate upwelling infrared radiation was modified to operate efficiently on the STAR-100. The modified software processes specific test cases significantly faster than the initial STAR-100 code. For example, a midlatitude summer atmospheric model is executed in less than 2% of the time originally required on the STAR-100. Furthermore, the optimized program performs extra operations to save the calculated absorption coefficients. Some of the advantages and pitfalls of virtual memory and vector processing are discussed along with strategies used to avoid loss of accuracy and computing power. Results from the vectorized code, in terms of speed, cost, and relative error with respect to serial code solutions are encouraging.

Rarig, P. L.↗

Advanced development of double-injection, deep-impurity semiconductor switches

Deep-impurity, double-injection devices, commonly refered to as (DI) squared devices, represent a class of semiconductor switches possessing a very high degree of tolerance to electron and neutron irradiation and to elevated temperature operation. These properties have caused them to be considered as attractive candidates for space power applications. The design, fabrication, and testing of several varieties of (DI) squared devices intended for power switching are described. All of these designs were based upon gold-doped silicon material. Test results, along with results of computer simulations of device operation, other calculations based upon the assumed mode of operation of (DI) squared devices, and empirical information regarding power semiconductor device operation and limitations, have led to the conculsion that these devices are not well suited to high-power applications. When operated in power circuitry configurations, they exhibit high-power losses in both the off-state and on-state modes. These losses are caused by phenomena inherent to the physics and material of the devices and cannot be much reduced by device design optimizations. The (DI) squared technology may, however, find application in low-power functions such as sensing, logic, and memory, when tolerance to radiation and temperature are desirable (especially is device performance is improved by incorporation of deep-level impurities other than gold.

Hanes, M. H.↗