Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

Energy Usage in an Embedded Space Vision Application on a Tiled Architecture

The need for greater autonomy in platforms such as planetary rovers is driving rapidly to codes that far overwhelm the capabilities of conventional space-qualified single core processors to run them in real-time. However, a new generation of potentially space-qualified 2D "tiled" multi-core microprocessor chips is emerging with significant performance potential. Leveraging such inherently parallel hardware for space platforms requires consideration of both time and power limitations - the latter of which is not normally done in conventional parallel computing. This paper takes one such application, Rockster, and analyzes it for energy usage when ported to a multi-core tiled chip such as may come from the Maestro program. The results demonstrate not only the criticality of memory and interconnect in the energy of real-time parallel codes, but also the effects of possible "energy-aware" changes in partitioning and algorithm design.

multi-core processors↗

Applications of the massively parallel machine, the MasPar MP-1, to Earth sciences

The computational workload of upcoming NASA science missions, especially the ground data processing for the Earth Observing System, is projected to be quite large (in the 50 to 100 gigaFLOPS range) and corespondingly very expensive to perform using conventional supercomputer systems. High performance, general purpose massively parallel computer systems such as the MasPar MP-1 are being investigated by NASA as a more cost effective alternative. Massively parallel systems are targeted for accelerated development and maturation by NASA's upcoming five-year High Performance Computing and Communications Program. A summary of the broad range of applications currently running on the MP-1 at NASA/Goddard are presented in this paper along with descriptions of the parallel algorithmic techniques employed in five applications that have bearing on Earth sciences.

Fischer, James R.↗

Innovative Language-Based & Object-Oriented Structured AMR Using Fortran 90 and OpenMP

Parallel adaptive mesh refinement (AMR) is an important numerical technique that leads to the efficient solution of many physical and engineering problems. In this paper, we describe how AMR programing can be performed in an object-oreinted way using the modern aspects of Fortran 90 combined with the parallelization features of OpenMP.

software↗

PROGRESS REPORT ON THE DEVELOPMENT OF PROTECTED CONSTRUCTION FOR HYPERSONIC VEHICLES

The structural problems of re-entry occasioned by aerodynamic heating are now generally well known. Typical of these heating problems is the environment experienced by the manned lifting vehicle re-entering from a low altitude orbit. Figure 1 shows typical curves of lower surface temperature as a function of time for a wing loading of 25 Ibs/ft(exp 2), and a lift-drag ratio of 2.5. The two curves cover a practical range of re-entry angles, the lower curve representing an ideal re-entry path with zero dive angle and the upper curve assuming a 2° error in the re-entry angle. Maneuvers for course change or correction are also added to the upper curve. Significant factors from this curve are the relationship between the maximum temperatures and the capabilities of available metallic material, and also the long re-entry time and its effect on total heat load. If a very shallow re-entry is made, maximum lower surface temperatures reach 2000°F which is just within the range of conventional superalloys. To accommodate practical re-entry angles and maneuvers, however, the temperature capability must reach about 2500°F which requires refractory metals. The re-entry time may be as high as 100 minutes, which gives total heat loads of approximately 40,000 BTU/ft(exp 2). A heat load of this magnitude, with equilibrium surface temperatures of the values shown, suggests that a lighter airframe can be constructed by dissipating the heat by radiation from the surface, rather than by absorbing it with a heat sink, or surface cooling, or ablation. Air Force programs to provide airframes for this type of environment have involved the parallel development of a number of different structural concepts. The development to be discussed here was carried out by Bell Aerosystems Company for the Fabrication Branch, Manufacturing Technology Laboratory, Directorate of Materials and Processes, Aeronautical Systems' Division, Wright-Patterson Air Force Base, Ohio, under Contract AF33(600)-40100 (Double-Wall Cooled Structure).

Reentry vehicle↗

Thrust imbalance of the Space Shuttle solid rocket motors

The Monte Carlo statistical analysis of thrust imbalance is applied to both the Titan IIIC and the Space Shuttle solid rocket motors (SRMs) firing in parallel, and results are compared with those obtained from the Space Shuttle program. The test results are examined in three phases: (1) pairs of SRMs selected from static tests of the four developmental motors (DMs 1 through 4); (2) pairs of SRMs selected from static tests of the three quality assurance motors (QMs 1 through 3); (3) SRMs on the first flight test vehicle (STS-1A and STS-1B). The simplified internal ballistic model utilized for computing thrust from head-end pressure measurements on flight tests is shown to agree closely with measured thrust data. Inaccuracies in thrust imbalance evaluation are explained by possible flight test instrumentation errors.

Foster, W. A., Jr.↗

Application of the p-version of the finite-element method to global-local problems

A brief survey is given of some recent developments in finite-element analysis technology which bear upon the three main research areas under consideration in this workshop: (1) analysis methods; (2) software testing and quality assurance; and (3) parallel processing. The variational principle incorporated in a finite-element computer program, together with a particular set of input data, determines the exact solution corresponding to that input data. Most finite-element analysis computer programs are based on the principle of virtual work. In the following, researchers consider only programs based on the principle of virtual work and denote the exact displacement vector field corresponding to some specific set of input data by vector u(EX). The exact solution vector u(EX) is independent of the design of the mesh or the choice of elements. Except for very simple problems, or specially constructed test problems, vector u(EX) is not known. Researchers perform a finite-element analysis (or any other numerical analysis) because they wish to make conclusions concerning the response of a physical system to certain imposed conditions, as if vector u(EX) were known.

Szabo, Barna A.↗

The PARTY parallel runtime system

In the present automated system for the organization of the data and computational operations entailed by parallel problems, in ways that optimize multiprocessor performance, general heuristics for partitioning program data and control are implemented by capturing and manipulating representations of a computation at run time. These heuristics are directed toward the dynamic identification and allocation of concurrent work in computations with irregular computational patterns. An optimized static-workload partitioning is computed for such repetitive-computation pattern problems as the iterative ones employed in scientific computation.

Saltz, J. H.↗

Load Variation Influences on Joint Work During Squat Exercise in Reduced Gravity

Resistance exercises that load the axial skeleton, such as the parallel squat, are incorporated as a critical component of a space exercise program designed to maximize the stimuli for bone remodeling and muscle loading. Astronauts on the International Space Station perform regular resistance exercise using the Advanced Resistive Exercise Device (ARED). Squat exercises on Earth entail moving a portion of the body weight plus the added bar load, whereas in microgravity the body weight is 0, so all load must be applied via the bar. Crewmembers exercising in microgravity currently add approx.70% of their body weight to the bar load as compensation for the absence of the body weight. This level of body weight replacement (BWR) was determined by crewmember feedback and personal experience without any quantitative data. The purpose of this evaluation was to utilize computational simulation to determine the appropriate level of BWR in microgravity necessary to replicate lower extremity joint work during squat exercise in normal gravity based on joint work. We hypothesized that joint work would be positively related to BWR load.

DeWitt, John K.↗

Real-Time Cognitive Computing Architecture for Data Fusion in a Dynamic Environment

A novel cognitive computing architecture is conceptualized for processing multiple channels of multi-modal sensory data streams simultaneously, and fusing the information in real time to generate intelligent reaction sequences. This unique architecture is capable of assimilating parallel data streams that could be analog, digital, synchronous/asynchronous, and could be programmed to act as a knowledge synthesizer and/or an "intelligent perception" processor. In this architecture, the bio-inspired models of visual pathway and olfactory receptor processing are combined as processing components, to achieve the composite function of "searching for a source of food while avoiding the predator." The architecture is particularly suited for scene analysis from visual data and odorant.

Duong, Tuan A.↗

Design and fabrication of a stereoscopic rear-viewing endoscopic tool (MARVEL)

The use of minimally invasive neurosurgical techniques has experienced a growing interest among patients and surgeons alike due their the numerous advantages over more traditional operations. Current methods employ the use either rigid or articulating endoscopes, both of which are inserted through a small diameter opening in the skull near the intended surgical region. Although these devices aid the surgeon in viewing the operating site, they are severely limited by their inability to provide dimensional awareness to the user. The Multi-Angle Rear-Viewing Endoscopic TooL (MARVEL) has been designed such that the compact form of existing endoscopes is maintained while also providing the user with an enhanced 3-dimensional stereo view. The design of the alpha prototype for MARVEL was completed over the course of three months. In parallel with development, the device's optics were characterized and several software programs were generated. Once assembled, MARVEL will hold many advantages over existing endoscopic tools. It is expected that the device will greatly outperform its traditional counterparts by both increasing safety and decreasing the duration of surgical procedures. After completion, the device will be demonstrated to the project sponsors at the Skull Base Institute where it will undergo full evaluation and a feasibility analysis.

Strongrich, Andrew↗

Design of a real-time wind turbine simulator using a custom parallel architecture

The design of a new parallel-processing digital simulator is described. The new simulator has been developed specifically for analysis of wind energy systems in real time. The new processor has been named: the Wind Energy System Time-domain simulator, version 3 (WEST-3). Like previous WEST versions, WEST-3 performs many computations in parallel. The modules in WEST-3 are pure digital processors, however. These digital processors can be programmed individually and operated in concert to achieve real-time simulation of wind turbine systems. Because of this programmability, WEST-3 is very much more flexible and general than its two predecessors. The design features of WEST-3 are described to show how the system produces high-speed solutions of nonlinear time-domain equations. WEST-3 has two very fast Computational Units (CU's) that use minicomputer technology plus special architectural features that make them many times faster than a microcomputer. These CU's are needed to perform the complex computations associated with the wind turbine rotor system in real time. The parallel architecture of the CU causes several tasks to be done in each cycle, including an IO operation and the combination of a multiply, add, and store. The WEST-3 simulator can be expanded at any time for additional computational power. This is possible because the CU's interfaced to each other and to other portions of the simulation using special serial buses. These buses can be 'patched' together in essentially any configuration (in a manner very similar to the programming methods used in analog computation) to balance the input/ output requirements. CU's can be added in any number to share a given computational load. This flexible bus feature is very different from many other parallel processors which usually have a throughput limit because of rigid bus architecture.

Hoffman, John A.↗

The Necessity of Functional Analysis for Space Exploration Programs

As NASA moves toward expanded commercial spaceflight within its human exploration capability, there is increased emphasis on how to allocate responsibilities between government and commercial organizations to achieve coordinated program objectives. The practice of program-level functional analysis offers an opportunity for improved understanding of collaborative functions among heterogeneous partners. Functional analysis is contrasted with the physical analysis more commonly done at the program level, and is shown to provide theoretical performance, risk, and safety advantages beneficial to a government-commercial partnership. Performance advantages include faster convergence to acceptable system solutions; discovery of superior solutions with higher commonality, greater simplicity and greater parallelism by substituting functional for physical redundancy to achieve robustness and safety goals; and greater organizational cohesion around program objectives. Risk advantages include avoidance of rework by revelation of some kinds of architectural and contractual mismatches before systems are specified, designed, constructed, or integrated; avoidance of cost and schedule growth by more complete and precise specifications of cost and schedule estimates; and higher likelihood of successful integration on the first try. Safety advantages include effective delineation of must-work and must-not-work functions for integrated hazard analysis, the ability to formally demonstrate completeness of safety analyses, and provably correct logic for certification of flight readiness. The key mechanism for realizing these benefits is the development of an inter-functional architecture at the program level, which reveals relationships between top-level system requirements that would otherwise be invisible using only a physical architecture. This paper describes the advantages and pitfalls of functional analysis as a means of coordinating the actions of large heterogeneous organizations for space exploration programs.

program management↗

Reducing neural network training time with parallel processing

Obtaining optimal solutions for engineering design problems is often expensive because the process typically requires numerous iterations involving analysis and optimization programs. Previous research has shown that a near optimum solution can be obtained in less time by simulating a slow, expensive analysis with a fast, inexpensive neural network. A new approach has been developed to further reduce this time. This approach decomposes a large neural network into many smaller neural networks that can be trained in parallel. Guidelines are developed to avoid some of the pitfalls when training smaller neural networks in parallel. These guidelines allow the engineer: to determine the number of nodes on the hidden layer of the smaller neural networks; to choose the initial training weights; and to select a network configuration that will capture the interactions among the smaller neural networks. This paper presents results describing how these guidelines are developed.

Rogers, James L., Jr.↗

Simulation/Emulation Techniques: Compressing Schedules With Parallel (HW/SW) Development

NASA has always been in the business of balancing new technologies and techniques to achieve human space travel objectives. NASA's Kedalion engineering analysis lab has been validating and using many contemporary avionics HW/SW development and integration techniques, which represent new paradigms to NASA's heritage culture. Kedalion has validated many of the Orion HW/SW engineering techniques borrowed from the adjacent commercial aircraft avionics solution space, inserting new techniques and skills into the Multi - Purpose Crew Vehicle (MPCV) Orion program. Using contemporary agile techniques, Commercial-off-the-shelf (COTS) products, early rapid prototyping, in-house expertise and tools, and extensive use of simulators and emulators, NASA has achieved cost effective paradigms that are currently serving the Orion program effectively. Elements of long lead custom hardware on the Orion program have necessitated early use of simulators and emulators in advance of deliverable hardware to achieve parallel design and development on a compressed schedule.

Mangieri, Mark L.↗

Regenerative fuel cell energy storage system for a low earth orbit space station

A study was conducted to define characteristics of a Regenerative Fuel Cell System (RFCS) for low earth orbit Space Station missions. The RFCS's were defined and characterized based on both an alkaline electrolyte fuel cell integrated with an alkaline electrolyte water electrolyzer and an alkaline electrolyte fuel cell integrated with an acid solid polymer electrolyte (SPE) water electrolyzer. The study defined the operating characteristics of the systems including system weight, volume, and efficiency. A maintenance philosophy was defined and the implications of system reliability requirements and modularization were determined. Finally, an Engineering Model System was defined and a program to develop and demonstrate the EMS and pacing technology items that should be developed in parallel with the EMS were identified. The specific weight of an optimized RFCS operating at 140 F was defined as a function of system efficiency for a range of module sizes. An EMS operating at a nominal temperature of 180 F and capable of delivery of 10 kW at an overall efficiency of 55.4 percent is described. A program to develop the EMS is described including a technology development effort for pacing technology items.

Martin, R. E.↗

Annual Research Briefs, 2002: Center for Turbulence Research

Turbulent combustion remains the largest component of the CTR's core program. This program and several related activities at CTR are supported by NASA's Ultra Efficient Engine Technology Program. It is also intimately connected with the Department of Energy's ASCI program at Stanford which develops the technology for numerical simulation of realistic aircraft engines using state of the art massively parallel computers. In combustion modeling the attention has been directed to the modeling of higher levels of complexity such as spray dynamics, radiation and soot formation. Major aircraft engine manufacturers have shown considerable interest in this program; in particular, a significant active collaboration exists between CTR and the Pratt & Whitney Corporation. CTR's combustion program is essentially based on the large-eddy simulation technique, LES, which is actively being pursued at CTR for this and many other applications. Important accomplishments in LES included simulations with three-dimensional filters which result in grid independent calculations (that is why we call it "true" LES), and the development of the methodology for integration of LES and Reynolds Averaged computations. Optimization techniques are being studied and used for the important problem of wall boundary conditions for LES as well as for optimal shape design for aeroacoustic and aerodynamic performance gains.

ULTRA EFFICIENT ENGINE TECHNOLOGY↗

Data Acquisition System for Multi-Frequency Radar Flight Operations Preparation

A three-channel data acquisition system was developed for the NASA Multi-Frequency Radar (MFR) system. The system is based on a commercial-off-the-shelf (COTS) industrial PC (personal computer) and two dual-channel 14-bit digital receiver cards. The decimated complex envelope representations of the three radar signals are passed to the host PC via the PCI bus, and then processed in parallel by multiple cores of the PC CPU (central processing unit). The innovation is this parallelization of the radar data processing using multiple cores of a standard COTS multi-core CPU. The data processing portion of the data acquisition software was built using autonomous program modules or threads, which can run simultaneously on different cores. A master program module calculates the optimal number of processing threads, launches them, and continually supplies each with data. The benefit of this new parallel software architecture is that COTS PCs can be used to implement increasingly complex processing algorithms on an increasing number of radar range gates and data rates. As new PCs become available with higher numbers of CPU cores, the software will automatically utilize the additional computational capacity.

Leachman, Jonathan↗

Implementing Access to Data Distributed on Many Processors

A reference architecture is defined for an object-oriented implementation of domains, arrays, and distributions written in the programming language Chapel. This technology primarily addresses domains that contain arrays that have regular index sets with the low-level implementation details being beyond the scope of this discussion. What is defined is a complete set of object-oriented operators that allows one to perform data distributions for domain arrays involving regular arithmetic index sets. What is unique is that these operators allow for the arbitrary regions of the arrays to be fragmented and distributed across multiple processors with a single point of access giving the programmer the illusion that all the elements are collocated on a single processor. Today's massively parallel High Productivity Computing Systems (HPCS) are characterized by a modular structure, with a large number of processing and memory units connected by a high-speed network. Locality of access as well as load balancing are primary concerns in these systems that are typically used for high-performance scientific computation. Data distributions address these issues by providing a range of methods for spreading large data sets across the components of a system. Over the past two decades, many languages, systems, tools, and libraries have been developed for the support of distributions. Since the performance of data parallel applications is directly influenced by the distribution strategy, users often resort to low-level programming models that allow fine-tuning of the distribution aspects affecting performance, but, at the same time, are tedious and error-prone. This technology presents a reusable design of a data-distribution framework for data parallel high-performance applications. Distributions are a means to express locality in systems composed of large numbers of processor and memory components connected by a network. Since distributions have a great effect on the performance of applications, it is important that the distribution strategy is flexible, so its behavior can change depending on the needs of the application. At the same time, high productivity concerns require that the user be shielded from error-prone, tedious details such as communication and synchronization.

James, Mark↗