Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Reducing neural network training time with parallel processing

Obtaining optimal solutions for engineering design problems is often expensive because the process typically requires numerous iterations involving analysis and optimization programs. Previous research has shown that a near optimum solution can be obtained in less time by simulating a slow, expensive analysis with a fast, inexpensive neural network. A new approach has been developed to further reduce this time. This approach decomposes a large neural network into many smaller neural networks that can be trained in parallel. Guidelines are developed to avoid some of the pitfalls when training smaller neural networks in parallel. These guidelines allow the engineer: to determine the number of nodes on the hidden layer of the smaller neural networks; to choose the initial training weights; and to select a network configuration that will capture the interactions among the smaller neural networks. This paper presents results describing how these guidelines are developed.

Rogers, James L., Jr.↗

Simulation/Emulation Techniques: Compressing Schedules With Parallel (HW/SW) Development

NASA has always been in the business of balancing new technologies and techniques to achieve human space travel objectives. NASA's Kedalion engineering analysis lab has been validating and using many contemporary avionics HW/SW development and integration techniques, which represent new paradigms to NASA's heritage culture. Kedalion has validated many of the Orion HW/SW engineering techniques borrowed from the adjacent commercial aircraft avionics solution space, inserting new techniques and skills into the Multi - Purpose Crew Vehicle (MPCV) Orion program. Using contemporary agile techniques, Commercial-off-the-shelf (COTS) products, early rapid prototyping, in-house expertise and tools, and extensive use of simulators and emulators, NASA has achieved cost effective paradigms that are currently serving the Orion program effectively. Elements of long lead custom hardware on the Orion program have necessitated early use of simulators and emulators in advance of deliverable hardware to achieve parallel design and development on a compressed schedule.

Mangieri, Mark L.↗

Regenerative fuel cell energy storage system for a low earth orbit space station

A study was conducted to define characteristics of a Regenerative Fuel Cell System (RFCS) for low earth orbit Space Station missions. The RFCS's were defined and characterized based on both an alkaline electrolyte fuel cell integrated with an alkaline electrolyte water electrolyzer and an alkaline electrolyte fuel cell integrated with an acid solid polymer electrolyte (SPE) water electrolyzer. The study defined the operating characteristics of the systems including system weight, volume, and efficiency. A maintenance philosophy was defined and the implications of system reliability requirements and modularization were determined. Finally, an Engineering Model System was defined and a program to develop and demonstrate the EMS and pacing technology items that should be developed in parallel with the EMS were identified. The specific weight of an optimized RFCS operating at 140 F was defined as a function of system efficiency for a range of module sizes. An EMS operating at a nominal temperature of 180 F and capable of delivery of 10 kW at an overall efficiency of 55.4 percent is described. A program to develop the EMS is described including a technology development effort for pacing technology items.

Martin, R. E.↗

Annual Research Briefs, 2002: Center for Turbulence Research

Turbulent combustion remains the largest component of the CTR's core program. This program and several related activities at CTR are supported by NASA's Ultra Efficient Engine Technology Program. It is also intimately connected with the Department of Energy's ASCI program at Stanford which develops the technology for numerical simulation of realistic aircraft engines using state of the art massively parallel computers. In combustion modeling the attention has been directed to the modeling of higher levels of complexity such as spray dynamics, radiation and soot formation. Major aircraft engine manufacturers have shown considerable interest in this program; in particular, a significant active collaboration exists between CTR and the Pratt & Whitney Corporation. CTR's combustion program is essentially based on the large-eddy simulation technique, LES, which is actively being pursued at CTR for this and many other applications. Important accomplishments in LES included simulations with three-dimensional filters which result in grid independent calculations (that is why we call it "true" LES), and the development of the methodology for integration of LES and Reynolds Averaged computations. Optimization techniques are being studied and used for the important problem of wall boundary conditions for LES as well as for optimal shape design for aeroacoustic and aerodynamic performance gains.

ULTRA EFFICIENT ENGINE TECHNOLOGY↗

Data Acquisition System for Multi-Frequency Radar Flight Operations Preparation

A three-channel data acquisition system was developed for the NASA Multi-Frequency Radar (MFR) system. The system is based on a commercial-off-the-shelf (COTS) industrial PC (personal computer) and two dual-channel 14-bit digital receiver cards. The decimated complex envelope representations of the three radar signals are passed to the host PC via the PCI bus, and then processed in parallel by multiple cores of the PC CPU (central processing unit). The innovation is this parallelization of the radar data processing using multiple cores of a standard COTS multi-core CPU. The data processing portion of the data acquisition software was built using autonomous program modules or threads, which can run simultaneously on different cores. A master program module calculates the optimal number of processing threads, launches them, and continually supplies each with data. The benefit of this new parallel software architecture is that COTS PCs can be used to implement increasingly complex processing algorithms on an increasing number of radar range gates and data rates. As new PCs become available with higher numbers of CPU cores, the software will automatically utilize the additional computational capacity.

Leachman, Jonathan↗

Implementing Access to Data Distributed on Many Processors

A reference architecture is defined for an object-oriented implementation of domains, arrays, and distributions written in the programming language Chapel. This technology primarily addresses domains that contain arrays that have regular index sets with the low-level implementation details being beyond the scope of this discussion. What is defined is a complete set of object-oriented operators that allows one to perform data distributions for domain arrays involving regular arithmetic index sets. What is unique is that these operators allow for the arbitrary regions of the arrays to be fragmented and distributed across multiple processors with a single point of access giving the programmer the illusion that all the elements are collocated on a single processor. Today's massively parallel High Productivity Computing Systems (HPCS) are characterized by a modular structure, with a large number of processing and memory units connected by a high-speed network. Locality of access as well as load balancing are primary concerns in these systems that are typically used for high-performance scientific computation. Data distributions address these issues by providing a range of methods for spreading large data sets across the components of a system. Over the past two decades, many languages, systems, tools, and libraries have been developed for the support of distributions. Since the performance of data parallel applications is directly influenced by the distribution strategy, users often resort to low-level programming models that allow fine-tuning of the distribution aspects affecting performance, but, at the same time, are tedious and error-prone. This technology presents a reusable design of a data-distribution framework for data parallel high-performance applications. Distributions are a means to express locality in systems composed of large numbers of processor and memory components connected by a network. Since distributions have a great effect on the performance of applications, it is important that the distribution strategy is flexible, so its behavior can change depending on the needs of the application. At the same time, high productivity concerns require that the user be shielded from error-prone, tedious details such as communication and synchronization.

James, Mark↗

Overview of NASA Gateway Lunar Dust Mitigation and Contamination Modeling and Analysis

The planned NASA Artemis campaign has several lunar surface missions in which the NASA Gateway is the waypoint between lunar orbit and the lunar surface for Human Lander System (HLS). Each of these missions is an opportunity for lunar dust to be introduced into the Gateway environment, post surface mission, potentially causing end of life (EOL) performance degradation due to lunar dust contamination of sensitive hardware and systems on the exterior of Gateway and Visiting Vehicles. The Gateway Systems Engineering and Integration (SE&I) and Induced Environments teams are addressing the challenge of lunar dust with a two-pronged approach. Characterization of the lunar dust induced environment around Gateway and contamination risk is accomplished with a comprehensive physics-based framework, the Gateway On-orbit Lunar Dust Modeling and Analysis Program (GOLDMAP), which is currently in development. Analysis results from GOLDMAP define Gateway-level induced environments requirements and are flowed to elements and subsystems. In parallel, a dust mitigation strategy is being developed with a focus on lunar dust protection, dust mitigation technologies, mitigation and testing guidance, and cross-program coordination and is informed by outputs from GOLDMAP analyses. Activities supporting both components include hardware susceptibility assessments and testing, and scientific experiments on lunar regolith.

Gateway↗

A real-time, dual processor simulation of the rotor system research aircraft

A real-time, man-in-the loop, simulation of the rotor system research aircraft (RSRA) was conducted. The unique feature of this simulation was that two digital computers were used in parallel to solve the equations of the RSRA mathematical model. The design, development, and implementation of the simulation are documented. Program validation was discussed, and examples of data recordings are given. This simulation provided an important research tool for the RSRA project in terms of safe and cost-effective design analysis. In addition, valuable knowledge concerning parallel processing and a powerful simulation hardware and software system was gained.

Mackie, D. B.↗

Function algorithms for MPP scientific subroutines, volume 1

Design documentation and user documentation for function algorithms for the Massively Parallel Processor (MPP) are presented. The contract specifies development of MPP assembler instructions to perform the following functions: natural logarithm; exponential (e to the x power); square root; sine; cosine; and arctangent. To fulfill the requirements of the contract, parallel array and solar implementations for these functions were developed on the PDP11/34 Program Development and Management Unit (PDMU) that is resident at the MPP testbed installation located at the NASA Goddard facility.

Gouch, J. G.↗

A direct-execution parallel architecture for the Advanced Continuous Simulation Language (ACSL)

A direct-execution parallel architecture for the Advanced Continuous Simulation Language (ACSL) is presented which overcomes the traditional disadvantages of simulations executed on a digital computer. The incorporation of parallel processing allows the mapping of simulations into a digital computer to be done in the same inherently parallel manner as they are currently mapped onto an analog computer. The direct-execution format maximizes the efficiency of the executed code since the need for a high level language compiler is eliminated. Resolution is greatly increased over that which is available with an analog computer without the sacrifice in execution speed normally expected with digitial computer simulations. Although this report covers all aspects of the new architecture, key emphasis is placed on the processing element configuration and the microprogramming of the ACLS constructs. The execution times for all ACLS constructs are computed using a model of a processing element based on the AMD 29000 CPU and the AMD 29027 FPU. The increase in execution speed provided by parallel processing is exemplified by comparing the derived execution times of two ACSL programs with the execution times for the same programs executed on a similar sequential architecture.

Carroll, Chester C.↗

Chemical calculations on Cray computers

The influence of recent developments in supercomputing on computational chemistry is discussed with particular reference to Cray computers and their pipelined vector/limited parallel architectures. After reviewing Cray hardware and software the performance of different elementary program structures are examined, and effective methods for improving program performance are outlined. The computational strategies appropriate for obtaining optimum performance in applications to quantum chemistry and dynamics are discussed. Finally, some discussion is given of new developments and future hardware and software improvements.

Taylor, Peter R.↗

Report from ionospheric science

The general strategy to advance knowledge of the ionospheric component of the solar terrestrial system should consist of a three pronged attack on the problem. Ionospheric models should be refined by utilization of existing and new data bases. The data generated in the future should emphasize spatial and temporal gradients and their relation to other events in the solar terrestrial system. In parallel with the improvement in modeling, it will be necessary to initiate a program of advanced instrument development. In particular, emphasis should be placed on the area of improved imaging techniques. The third general activity to be supported should be active experiments related to a better understanding of the basic physics of interactions occurring in the ionospheric environment. These strategies are briefly discussed.

Raitt, W. J.↗

Chemical calculations on Cray computers

The influence of recent developments in supercomputing on computational chemistry is discussed with particular reference to Cray computers and their pipelined vector/limited parallel architectures. After reviewing Cray hardware and software the performance of different elementary program structures are examined, and effective methods for improving program performance are outlined. The computational strategies appropriate for obtaining optimum performance in applications to quantum chemistry and dynamics are discussed. Finally, some discussion is given of new developments and future hardware and software improvements.

Taylor, Peter R.↗

Discovery: Near-Earth Asteroid Rendezvous (NEAR)

The work carried out under this grant consisted of two parallel studies aimed at defining candidate missions for the initiation of the Discovery Program being considered by NASA's Solar System Exploration Division. The main study considered a Discover-class mission to a Near Earth Asteroid (NEA); the companion study considered a small telescope in Earth-orbit dedicated to ultra violet studies of solar system bodies. The results of these studies are summarized in two reports which are attached (Appendix 1 and Appendix 2).

Veverka, Joseph↗

Numerical experiments with flows of elongated granules

Theory and numerical results are given for a program simulating two dimensional granular flow (1) between two infinite, counter-moving, parallel, roughened walls, and (2) for an infinitely wide slider. Each granule is simulated by a central repulsive force field ratcheted with force restitution factor to introduce dissipation. Transmission of angular momentum between particles occurs via Coulomb friction. The effect of granular hardness is explored. Gaps from 7 to 28 particle diameters are investigated, with solid fractions ranging from 0.2 to 0.9. Among features observed are: slip flow at boundaries, coagulation at high densities, and gross fluctuation in surface stress. A videotape has been prepared to demonstrate the foregoing effects.

Elrod, Harold G.↗

Computationally efficient multibody simulations

Computationally efficient approaches to the solution of the dynamics of multibody systems are presented in this work. The computational efficiency is derived from both the algorithmic and implementational standpoint. Order(n) approaches provide a new formulation of the equations of motion eliminating the assembly and numerical inversion of a system mass matrix as required by conventional algorithms. Computational efficiency is also gained in the implementation phase by the symbolic processing and parallel implementation of these equations. Comparison of this algorithm with existing multibody simulation programs illustrates the increased computational efficiency.

Ramakrishnan, Jayant↗

Navier-Stokes Aerodynamic Simulation of the V-22 Osprey on the Intel Paragon MPP

The paper will describe the Development of a general three-dimensional multiple grid zone Navier-Stokes flowfield simulation program (ENS3D-MPP) designed for efficient execution on the Intel Paragon Massively Parallel Processor (MPP) supercomputer, and the subsequent application of this method to the prediction of the viscous flowfield about the V-22 Osprey tiltrotor vehicle. The flowfield simulation code solves the thin Layer or full Navier-Stoke's equation - for viscous flow modeling, or the Euler equations for inviscid flow modeling on a structured multi-zone mesh. In the present paper only viscous simulations will be shown. The governing difference equations are solved using a time marching implicit approximate factorization method with either TVD upwind or central differencing used for the convective terms and central differencing used for the viscous diffusion terms. Steady state or Lime accurate solutions can be calculated. The present paper will focus on steady state applications, although time accurate solution analysis is the ultimate goal of this effort. Laminar viscosity is calculated using Sutherland's law and the Baldwin-Lomax two layer algebraic turbulence model is used to compute the eddy viscosity. The Simulation method uses an arbitrary block, curvilinear grid topology. An automatic grid adaption scheme is incorporated which concentrates grid points in high density gradient regions. A variety of user-specified boundary conditions are available. This paper will present the application of the scalable and superscalable versions to the steady state viscous flow analysis of the V-22 Osprey using a multiple zone global mesh. The mesh consists of a series of sheared cartesian grid blocks with polar grids embedded within to better simulate the wing tip mounted nacelle. MPP solutions will be shown in comparison to equivalent Cray C-90 results and also in comparison to experimental data. Discussions on meshing considerations, wall clock execution time, load balancing, and scalability will be provided.

Vadyak, Joseph↗

Computation of Earth Science Products on Spaceborne Platforms

Spaceborne sensors like NASA's Hyperion hyperspectral imager generate huge data volumes, and several near-term trends indicate that data volumes will only increase. Next-generation hyperspectral missions, such as NASA's Hyperspectral Infrared Imager (HyspIRI), will operate at higher duty cycles and higher data rates, and their users will expect products to be generated from the data in near real time [1]. Barring a sudden advance in satellite downlink capacity, these trends point to a need to process data and generate products onboard the spacecraft. Rather than downlink an entire hyperspectral image cube, onboard processing enables satellites to downlink partial or completed scientific data products, which are often one to two orders of magnitude smaller than the original image. In addition, a satellite with onboard data processing resources and direct broadcast transmission equipment could send data products directly to first responders, research scientists or other users on the ground. Next-generation space-capable data processors will have a combination of reconfigurable gate arrays, digital signal processors and general-purpose CPUs. Correctly programmed and configured, these resources are sufficient to run sophisticated data analysis programs, including hyperspectral image processing algorithms that commonly run on desktop computers [2]. This paper describes how we implemented one such program, the HSEG hierarchical image segmentation algorithm, software commonly used on desktop and parallel processors, on a hardware platform designed to mimic a next-generation space-capable data processor [3]. We also describe our approach to porting the algorithm to and optimizing it for the new platform, and determine the expected performance gains enabled by our design. This extended abstract will describe the HSEG algorithm and hardware platform in greater detail, provide an analysis of the key function within the algorithm that required hardware acceleration, and describe our implementation of that function in hardware.

Fisher, Kevin↗