Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Measuring Low-Order Aberrations in a Segmented Telescope

The in-focus PSF optimizer (IPO) is an algorithm for use in monitoring and controlling the alignment of the segments of a segmented-mirror astronomical telescope. IPO is so named because it computes wave-front aberrations of the telescope from digitized pointspread functions (PSFs) measured in infocus images. Inasmuch as distant astronomical objects that behave optically as point sources can typically be seen in almost any astronomical image, the main benefit afforded by IPO may be to enable maintenance of mirror-segment alignments without detracting from valuable scientific-observation time. IPO evolved from prescription-retrieval type algorithms. Prescription retrieval uses in-focus and out-of-focus PSFs to infer the state of an imaging optical system. The state, in this context, refers to the positions, orientations, and low-order figure errors of the optical elements in the system. Both prescription- retrieval and IPO use an iterative, nonlinear, least-squares optimizer to compute the optimal state parameters such that a digital computer-generated model image matches the digitized image acquired from the real system. The difference between IPO and prescription- retrieval algorithms is that IPO is specifically designed to utilize infocus images only. Although the restriction to in-focus images limits IPO to calculating only the lowest-order wave front aberrations, it also causes the resulting computation to take much less time because fewer degrees of freedom are included in the optimization process. In the prescription retrieval software developed at JPL, the model images are generated using the ray-trace/physical optics program, MACOS. IPO, on the other hand, uses a linear sensitivity matrix to compute the exit-pupil wave front from the system parameters; the wave front is then converted into a complex pupil field, which is then propagated to the image plane via a fast Fourier transform. This approach is computationally faster and requires less computer memory than is needed for prescription retrieval.

Ohara, Catherine↗

Mechanisms of selective attention and space motion sickness

The neural mismatch theory of space motion sickness asserts that the central and peripheral autonomic sequelae of discordant sensory input arise from central integrative processes falling to reconcile patterns of incoming sensory information with existing memory. Stated differently, perceived novelty reaches a stress level as integrative mechanisms fail to return a sense of control to the individual in the new environment. Based on evidence summarized here, the severity of the neural mismatch may be dependent upon the relative amount of attention selectively afforded to each sensory input competing for control of behavior. Components of the limbic system may play important roles in match-mismatch operations, be therapeutically modulated by antimotion sickness drugs, and be optimally positioned to control autonomic output.

NASA Discipline Neuroscience↗

Algorithms for Automatic Alignment of Arrays

Aggregate data objects (such as arrays) are distributed across the processor memories when compiling a data-parallel language for a distributed-memory machine. The mapping determines the amount of communication needed to bring operands of parallel operations into alignment with each other. A common approach is to break the mapping into two stages: an alignment that maps all the objects to an abstract template, followed by a distribution that maps the template to the processors. This paper describes algorithms for solving the various facets of the alignment problem: axis and stride alignment, static and mobile offset alignment, and replication labeling. We show that optimal axis and stride alignment is NP-complete for general program graphs, and give a heuristic method that can explore the space of possible solutions in a number of ways. We show that some of these strategies can give better solutions than a simple greedy approach proposed earlier. We also show how local graph contractions can reduce the size of the problem significantly without changing the best solution. This allows more complex and effective heuristics to be used. We show how to model the static offset alignment problem using linear programming, and we show that loop-dependent mobile offset alignment is sometimes necessary for optimum performance. We describe an algorithm with for determining mobile alignments for objects within do loops. We also identify situations in which replicated alignment is either required by the program itself or can be used to improve performance. We describe an algorithm based on network flow that replicates objects so as to minimize the total amount of broadcast communication in replication.

Chatterjee, Siddhartha↗

PROCRU: A model for analyzing crew procedures in approach to landing

A model for analyzing crew procedures in approach to landing is developed. The model employs the information processing structure used in the optimal control model and in recent models for monitoring and failure detection. Mechanisms are added to this basic structure to model crew decision making in this multi task environment. Decisions are based on probability assessments and potential mission impact (or gain). Sub models for procedural activities are included. The model distinguishes among external visual, instrument visual, and auditory sources of information. The external visual scene perception models incorporate limitations in obtaining information. The auditory information channel contains a buffer to allow for storage in memory until that information can be processed.

Baron, S.↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Crowley, Kay↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Saltz, Joel↗

Portable Computer Technology (PCT) Research and Development Program Phase 2

The subject of this project report, focused on: (1) Design and development of two Advanced Portable Workstation 2 (APW 2) units. These units incorporate advanced technology features such as a low power Pentium processor, a high resolution color display, National Television Standards Committee (NTSC) video handling capabilities, a Personal Computer Memory Card International Association (PCMCIA) interface, and Small Computer System Interface (SCSI) and ethernet interfaces. (2) Use these units to integrate and demonstrate advanced wireless network and portable video capabilities. (3) Qualification of the APW 2 systems for use in specific experiments aboard the Mir Space Station. A major objective of the PCT Phase 2 program was to help guide future choices in computing platforms and techniques for meeting National Aeronautics and Space Administration (NASA) mission objectives. The focus being on the development of optimal configurations of computing hardware, software applications, and network technologies for use on NASA missions.

Castillo, Michael↗

Shape Memory Alloy Actuator Design: CASMART Collaborative Best Practices

Upon examination of shape memory alloy (SMA) actuation designs, there are many considerations and methodologies that are common to them all. A goal of CASMART's design working group is to compile the collective experiences of CASMART's member organizations into a single medium that engineers can then use to make the best decisions regarding SMA system design. In this paper, a review of recent work toward this goal is presented, spanning a wide range of design aspects including evaluation, properties, testing, modeling, alloy selection, fabrication, actuator processing, design optimization, controls, and system integration. We have documented each aspect, based on our collective experiences, so that the design engineer may access the tools and information needed to successfully design and develop SMA systems. Through comparison of several case studies, it is shown that there is not an obvious single, linear route a designer can adopt to navigate the path of concept to product. SMA engineering aspects will have different priorities and emphasis for different applications.

Benafan, Othmane↗

General-purpose interface bus for multiuser, multitasking computer system

The architecture of a multiuser, multitasking, virtual-memory computer system intended for the use by a medium-size research group is described. There are three central processing units (CPU) in the configuration, each with 16 MB memory, and two 474 MB hard disks attached. CPU 1 is designed for data analysis and contains an array processor for fast-Fourier transformations. In addition, CPU 1 shares display images viewed with the image processor. CPU 2 is designed for image analysis and display. CPU 3 is designed for data acquisition and contains 8 GPIB channels and an analog-to-digital conversion input/output interface with 16 channels. Up to 9 users can access the third CPU simultaneously for data acquisition. Focus is placed on the optimization of hardware interfaces and software, facilitating instrument control, data acquisition, and processing.

Generazio, Edward R.↗

Parallelization of NAS Benchmarks for Shared Memory Multiprocessors

This paper presents our experiences of parallelizing the sequential implementation of NAS benchmarks using compiler directives on SGI Origin2000 distributed shared memory (DSM) system. Porting existing applications to new high performance parallel and distributed computing platforms is a challenging task. Ideally, a user develops a sequential version of the application, leaving the task of porting to new generations of high performance computing systems to parallelization tools and compilers. Due to the simplicity of programming shared-memory multiprocessors, compiler developers have provided various facilities to allow the users to exploit parallelism. Native compilers on SGI Origin2000 support multiprocessing directives to allow users to exploit loop-level parallelism in their programs. Additionally, supporting tools can accomplish this process automatically and present the results of parallelization to the users. We experimented with these compiler directives and supporting tools by parallelizing sequential implementation of NAS benchmarks. Results reported in this paper indicate that with minimal effort, the performance gain is comparable with the hand-parallelized, carefully optimized, message-passing implementations of the same benchmarks.

Waheed, Abdul↗

Reusable, Extensible High-Level Data-Distribution Concept

A framework for high-level specification of data distributions in data-parallel application programs has been conceived. [As used here, distributions signifies means to express locality (more specifically, locations of specified pieces of data) in a computing system composed of many processor and memory components connected by a network.] Inasmuch as distributions exert a great effect on the performances of application programs, it is important that a distribution strategy be flexible, so that distributions can be adapted to the requirements of those programs. At the same time, for the sake of productivity in programming and execution, it is desirable that users be shielded from such error-prone, tedious details as those of communication and synchronization. As desired, the present framework enables a user to refine a distribution type and adjust it to optimize the performance of an application program and conceals, from the user, the low-level details of communication and synchronization. The framework provides for a reusable, extensible, data-distribution design, denoted the design pattern, that is independent of a concrete implementation. The design pattern abstracts over coding patterns that have been found to be commonly encountered in both manually and automatically generated distributed parallel programs. The following description of the present framework is necessarily oversimplified to fit within the space available for this article. Distributions are among the elements of a conceptual data-distribution machinery, some of the other elements being denoted domains, index sets, and data collections (see figure). Associated with each domain is one index set and one distribution. A distribution class interface (where "class" is used in the object-oriented-programming sense) includes operations that enable specification of the mapping of an index to a unit of locality. Thus, "Map(Index)" specifies a unit, while "LocalLayout(Index)" specifies the local address within that unit. The distribution class can be extended to enable specification of commonly used distributions or novel user-defined distributions. A data collection can be defined over a domain. The term "data collection" in this context signifies, more specifically, an abstraction of mappings from index sets to variables. Since the index set is distributed, the addresses of the variables are also distributed.

James, Mark↗

Developing a Hydrological Monitoring and Sub-Seasonal to Seasonal Forecasting System for South and Southeast Asian River Basins

South and Southeast Asia is subject to significant hydrometeorological extremes, including drought. Under rising temperatures, growing populations, and an apparent weakening of the South Asian monsoon in recent decades, concerns regarding drought and its potential impacts on water and food security are on the rise. Reliable sub-seasonal to seasonal (S2S) hydrological forecasts could, in principle, help governments and international organizations to better assess risk and act in the face of an oncoming drought. Here, we leverage recent improvements in S2S meteorological forecasts and the growing power of Earth Observations to provide more accurate monitoring of hydrological states for forecast initialization. Information from both sources is merged in a South and Southeast Asia sub-seasonal to seasonal hydrological forecasting system (SAHFS-S2S), developed collaboratively with the NASA SERVIR program and end-users across the region. This system applies the Noah-MultiParameterization (NoahMP) Land Surface Model (LSM) in the NASA Land Information System (LIS), driven by downscaled meteorological fields from the Global Data Assimilation System (GDAS) and Climate Hazards InfraRed Precipitation products (CHIRP and CHIRPS) to optimize initial conditions. The NASA Goddard Earth Observing System Model - sub-seasonal to seasonal (GEOS-S2S) forecasts, downscaled using the National Center for Atmospheric Research (NCAR) General Analog Regression Downscaling (GARD) tool and quantile mapping, are then applied to drive 5-km resolution hydrological forecasts to a 9-month forecast time horizon. Results show that the skillful predictions of root zone soil moisture can be made one to two months in advance for forecasts initialized in rainy seasons and up to 8 months when initialized in dry seasons. The memory of accurate initial conditions can positively contribute to forecast skills throughout the entire 9-month prediction period in areas with limited precipitation. This SAHFS-S2S has been operationalized at the International Centre for Integrated Mountain Development (ICIMOD) to support drought monitoring and warning needs in the region.

Yifan Zhou↗

Practical implementation of an accurate method for multilevel design sensitivity analysis

Solution techniques for handling large scale engineering optimization problems are reviewed. Potentials for practical applications as well as their limited capabilities are discussed. A new solution algorithm for design sensitivity is proposed. The algorithm is based upon the multilevel substructuring concept to be coupled with the adjoint method of sensitivity analysis. There are no approximations involved in the present algorithm except the usual approximations introduced due to the discretization of the finite element model. Results from the six- and thirty-bar planar truss problems show that the proposed multilevel scheme for sensitivity analysis is more effective (in terms of computer incore memory and the total CPU time) than a conventional (one level) scheme even on small problems. The new algorithm is expected to perform better for larger problems and its applications on the new generation of computer hardwares with 'parallel processing' capability is very promising.

Nguyen, Duc T.↗

Trends in modern system theory

The topics considered are related to linear control system design, adaptive control, failure detection, control under failure, system reliability, and large-scale systems and decentralized control. It is pointed out that the design of a linear feedback control system which regulates a process about a desirable set point or steady-state condition in the presence of disturbances is a very important problem. The linearized dynamics of the process are used for design purposes. The typical linear-quadratic design involving the solution of the optimal control problem of a linear time-invariant system with respect to a quadratic performance criterion is considered along with gain reduction theorems and the multivariable phase margin theorem. The stumbling block in many adaptive design methodologies is associated with the amount of real time computation which is necessary. Attention is also given to the desperate need to develop good theories for large-scale systems, the beginning of a microprocessor revolution, the translation of the Wiener-Hopf theory into the time domain, and advances made in dynamic team theory, dynamic stochastic games, and finite memory stochastic control.

Athans, M.↗

Multiphase complete exchange on a circuit switched hypercube

On a distributed memory parallel computer, the complete exchange (all-to-all personalized) communication pattern requires each of n processors to send a different block of data to each of the remaining n - 1 processors. This pattern is at the heart of many important algorithms, most notably the matrix transpose. For a circuit switched hypercube of dimension d(n = 2(sup d)), two algorithms for achieving complete exchange are known. These are (1) the Standard Exchange approach that employs d transmissions of size 2(sup d-1) blocks each and is useful for small block sizes, and (2) the Optimal Circuit Switched algorithm that employs 2(sup d) - 1 transmissions of 1 block each and is best for large block sizes. A unified multiphase algorithm is described that includes these two algorithms as special cases. The complete exchange on a hypercube of dimension d and block size m is achieved by carrying out k partial exchange on subcubes of dimension d(sub i) Sigma(sup k)(sub i=1) d(sub i) = d and effective block size m(sub i) = m2(sup d-di). When k = d and all d(sub i) = 1, this corresponds to algorithm (1) above. For the case of k = 1 and d(sub i) = d, this becomes the circuit switched algorithm (2). Changing the subcube dimensions d, varies the effective block size and permits a compromise between the data permutation and block transmission overhead of (1) and the startup overhead of (2). For a hypercube of dimension d, the number of possible combinations of subcubes is p(d), the number of partitions of the integer d. This is an exponential but very slowly growing function and it is feasible over these partitions to discover the best combination for a given message size. The approach was analyzed for, and implemented on, the Intel iPSC-860 circuit switched hypercube. Measurements show good agreement with predictions and demonstrate that the multiphase approach can substantially improve performance for block sizes in the 0 to 160 byte range. This range, which corresponds to 0 to 40 floating point numbers per processor, is commonly encountered in practical numeric applications. The multiphase technique is applicable to all circuit-switched hypercubes that use the common e-cube routing strategy.

Bokhari, Shahid H.↗

NASA Tech Briefs, November 1995

The contents include: 1) Mission Accomplished; 2) Resource Report: Marshall Space Flight Center; 3) NASA 1995 Software of the Year Award; 4) Microbolometers Based on Epitaxial YBa2Cu3O(sub 7-x) Thin Films; 5) Garnet Random-Access Memory; 6) Fabrication of SNS Weak Links on SOS Substrates; 7) High-Voltage MOSFET Switching Circuit; 8) Asymmetric Switching for a PWM H-Bridge Power Circuit; 9) Better Ohmic Contacts for InP Semiconductor Devices; 10) Low-Bandgap Thermovoltaic Materials and Devices; 11) Digital Frequency-Differencing Circuit; 12) Imaging Magnetometer; 13) Computer-Assisted Monitoring of a Complex System; 14) Buffered Telemetry Demodulator; 15) Compact Multifunction Inspection Head; 16) Optical Detection of Fractures in Ceramic Diaphragms; 17) Eddy-Current Detection of Cracks in Reinforced Carbon/Carbon; 18) Apparent Thermal Conductivity of Multilayer Insulation; 19) Optimizing Misch-Metal Compositions in Metal Hydride Anodes; 20) Device for Sampling Surface Contamination; 21) Probabilistic Failure Assessment for Fatigue; 22) Probabilistic Fatigue and Flaw-Propagation Analysis; 23) Windows Program for Driving the TDU-850 Printer; 24) Subband/Transform MATLAB Functions for Processing Images; 25) Computing Equilibrium Chemical Compositions; 26) Program Processes Thermocouple Readings; 27) ICAN-Second-Generation Integrated Composite Analyzer; 28) Integrated Composite Analyzer with Damping Capabilities; 29) Computing Efficiency of Transfer of Microwave Power; 30) Program Calculates Power Demands of Electronic Designs; 31) Cost-Estimation Program; 32) Program Estimates Areas Required by Electronic Designs; 33) Program to Balance Mapped Turbopump Assemblies; 34) BiblioTech; 35) Controlling Mirror Tilt With a Bimorph Actuator; 36) Burst-Disk Device Simulates Effect of Pyrotechnic Device; 37) Bearing-Mounting Concept Accommodates Thermal Expansion; 38) Parallel-Plate Acoustic Absorbers for Hot Environments; 39) Adjustable-Length Strut Withstands Large Cyclic Loads; 40) Tool Indicates Contact Angles in Bearing Raceways; 41) Gravity Slides With Magnetic Braking; 42) High-Torque, Lightweight, Pneumatically Driven Wrench for Small Spaces; 43) Device for Testing Compatibility of an O-Ring; 44) Magnetic Heat Pump Containing Flow Diverters; 45) Variable-Tilt Helicopter Rotor Mast; 46) "Beach-Ball" Robotic Rovers; 47) Apparatus Would Measure Temperatures of Ball Bearings; 48) Flexible Borescope for Inspecting Ducts; 49) Texturing Copper To Reduce Secondary Emission of Electrons; 50) Automated Laser Cutting in Three Dimensions; 51) Algorithm Helps Monitor Engine Operation; 52) Flexible Revision of Data-Processing Communications; 53) Software for Managing the Use of Land; 54) Thermal Strap Increases Cryocooling Efficiency; 55) Reversible Nut With Engagement Indication; 56) Control Algorithms for Kinematically Redundant Manipulators; 57) Computed Hydrogen-Flow Splits in a Rocket Engine; 58) Pressure and Thermal Modeling of Rocket Launches; 59) Field of View of a Spacecraft Antenna: Analysis and Software; 60) Digital Controller for Laser-Beam-Steering Subsystem; 61) More About Beam-Steering Subsystem for Laser Communication; 62) Digital Controller for Laser-Beam-Steering Subsystem: Part 2; 63) Interface Circuit Board for Space-Shuttle Communications; 64) Automated Planning of Spacecraft Telecommunications; 65) Artifacts of Spectral Analysis of Instrument Readings; 66) Neural-Network Controller for Vibration Suppression; 67) Adaptive Finite-Element Computation in Fracture Mechanics; 68) Attitude Control for the Cassini Spacecraft; 69) Analytical Model for Fluid Dynamics in a Microgravity Environment; 70) Study of Rocket-Engine Joints Bonded by NVCU/NARloy-Z; 71) Improved Silicon Nitride for Advanced Heat Engines; 72) Parameters for Welding Aluminum/Lithium Alloys; 73) Lightweight Composite Intertank Structure; 74) Foil Patches Seal Small Vacuum Leaks; 75) Data Base on Cables and Connectors; 76) Effect of Clock Mode on Radiation Hardnessf an ADC; and 77) Fault-Tolerant Control for a Robotic Inspection System.

Source record↗

WNN 92; Proceedings of the 3rd Workshop on Neural Networks: Academic/Industrial/NASA/Defense, Auburn Univ., AL, Feb. 10-12, 1992 and South Shore Harbour, TX, Nov. 4-6, 1992

The present conference discusses such neural networks (NN) related topics as their current development status, NN architectures, NN learning rules, NN optimization methods, NN temporal models, NN control methods, NN pattern recognition systems and applications, biological and biomedical applications of NNs, VLSI design techniques for NNs, NN systems simulation, fuzzy logic, and genetic algorithms. Attention is given to missileborne integrated NNs, adaptive-mixture NNs, implementable learning rules, an NN simulator for travelling salesman problem solutions, similarity-based forecasting, NN control of hypersonic aircraft takeoff, NN control of the Space Shuttle Arm, an adaptive NN robot manipulator controller, a synthetic approach to digital filtering, NNs for speech analysis, adaptive spline networks, an anticipatory fuzzy logic controller, and encoding operations for fuzzy associative memories.

Padgett, Mary L.↗

Global Load Balancing with Parallel Mesh Adaption on Distributed-Memory Systems

Dynamic mesh adaption on unstructured grids is a powerful tool for efficiently computing unsteady problems to resolve solution features of interest. Unfortunately, this causes load imbalance among processors on a parallel machine. This paper describes the parallel implementation of a tetrahedral mesh adaption scheme and a new global load balancing method. A heuristic remapping algorithm is presented that assigns partitions to processors such that the redistribution cost is minimized. Results indicate that the parallel performance of the mesh adaption code depends on the nature of the adaption region and show a 35.5X speedup on 64 processors of an SP2 when 35% of the mesh is randomly adapted. For large-scale scientific computations, our load balancing strategy gives almost a sixfold reduction in solver execution times over non-balanced loads. Furthermore, our heuristic remapper yields processor assignments that are less than 3% off the optimal solutions but requires only 1% of the computational time.

Biswas, Rupak↗