Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗

Hardware Implementation of Lossless Adaptive and Scalable Hyperspectral Data Compression for Space

On-board lossless hyperspectral data compression reduces data volume in order to meet NASA and DoD limited downlink capabilities. The technique also improves signature extraction, object recognition and feature classification capabilities by providing exact reconstructed data on constrained downlink resources. At JPL a novel, adaptive and predictive technique for lossless compression of hyperspectral data was recently developed. This technique uses an adaptive filtering method and achieves a combination of low complexity and compression effectiveness that far exceeds state-of-the-art techniques currently in use. The JPL-developed 'Fast Lossless' algorithm requires no training data or other specific information about the nature of the spectral bands for a fixed instrument dynamic range. It is of low computational complexity and thus well-suited for implementation in hardware. A modified form of the algorithm that is better suited for data from pushbroom instruments is generally appropriate for flight implementation. A scalable field programmable gate array (FPGA) hardware implementation was developed. The FPGA implementation achieves a throughput performance of 58 Msamples/sec, which can be increased to over 100 Msamples/sec in a parallel implementation that uses twice the hardware resources This paper describes the hardware implementation of the 'Modified Fast Lossless' compression algorithm on an FPGA. The FPGA implementation targets the current state-of-the-art FPGAs (Xilinx Virtex IV and V families) and compresses one sample every clock cycle to provide a fast and practical real-time solution for space applications.

FPGA implementation↗

SAR processing on the MPP

The processing of synthetic aperture radar (SAR) signals using the massively parallel processor (MPP) is discussed. The fast Fourier transform convolution procedures employed in the algorithms are described. The MPP architecture comprises an array unit (ARU) which processes arrays of data; an array control unit which controls the operation of the ARU and performs scalar arithmetic; a program and data management unit which controls the flow of data; and a unique staging memory (SM) which buffers and permutes data. The ARU contains a 128 by 128 array of bit-serial processing elements (PE). Two-by-four surarrays of PE's are packaged in a custom VLSI HCMOS chip. The staging memory is a large multidimensional-access memory which buffers and permutes data flowing with the system. Efficient SAR processing is achieved via ARU communication paths and SM data manipulation. Real time processing capability can be realized via a multiple ARU, multiple SM configuration.

Batcher, K. E.↗

Government-to-government cooperation in space station development

A memoranda of understanding was recently signed between the United States (NASA) and three international Space Station partners - Canada, European Space Agency (ESA), and Japan. The international partners are performing parallel Phase B preliminary design studies, concurrent with the U.S., on their proposed elements/systems for possible integration and operation with the U.S. Space Station System complex. During the 21-month Space Station Phase B study, a large amount of technical interface data will have to be transferred between the U.S. and the international partners. Scheduled bilateral technical coordination meetings will also be held. The coordination and large number of interfaces required to integrate the international requirements into the Space Station require a clean interface management organizational structure and operation procedures to accomplish the integration task. The international coordination management organizational structure, management tools, and communications network are discussed including the proposed international elements/systems being studied by the international partners.

Nassiff, S. H.↗

A system for remote measurements of the wind stress over the ocean

The DISSTRESS system for remote measurements of the surface wind stress over the ocean from ships and buoys is described. It is fully digital, utilizing the inertial dissipation technique. Parallel processing allows anemometer data to be filtered in natural frequency space; that is, the filter cutoffs shift linearly with the mean wind speed of the data to be filtered. The construction of the digital Butterworth bandpass filters is presented in detail. The performance of the system is evaluated by analyzing the results from 28 days of operation during the Frontal Air-Sea Interaction Experiment. The mean wind speed is checked, the anemometer response function is established, and drag coefficients are compared to previous studies. The capability of the system is demonstrated by continuous time series of the friction velocity computed every 20 min. The conclusion is that the surface wind stress can be measured more reliably and accurately (20 percent) with this system than from anemometer wind speeds and a bulk formula.

Large, William G.↗

Applications of massively parallel computers in telemetry processing

Telemetry processing refers to the reconstruction of full resolution raw instrumentation data with artifacts, of space and ground recording and transmission, removed. Being the first processing phase of satellite data, this process is also referred to as level-zero processing. This study is aimed at investigating the use of massively parallel computing technology in providing level-zero processing to spaceflights that adhere to the recommendations of the Consultative Committee on Space Data Systems (CCSDS). The workload characteristics, of level-zero processing, are used to identify processing requirements in high-performance computing systems. An example of level-zero functions on a SIMD MPP, such as the MasPar, is discussed. The requirements in this paper are based in part on the Earth Observing System (EOS) Data and Operation System (EDOS).

El-Ghazawi, Tarek A.↗

An Atmospheric General Circulation Model with Chemistry for the CRAY T3E: Design, Performance Optimization and Coupling to an Ocean Model

The design, implementation and performance optimization on the CRAY T3E of an atmospheric general circulation model (AGCM) which includes the transport of, and chemical reactions among, an arbitrary number of constituents is reviewed. The parallel implementation is based on a two-dimensional (longitude and latitude) data domain decomposition. Initial optimization efforts centered on minimizing the impact of substantial static and weakly-dynamic load imbalances among processors through load redistribution schemes. Recent optimization efforts have centered on single-node optimization. Strategies employed include loop unrolling, both manually and through the compiler, the use of an optimized assembler-code library for special function calls, and restructuring of parts of the code to improve data locality. Data exchanges and synchronizations involved in coupling different data-distributed models can account for a significant fraction of the running time. Therefore, the required scattering and gathering of data must be optimized. In systems such as the T3E, there is much more aggregate bandwidth in the total system than in any particular processor. This suggests a distributed design. The design and implementation of a such distributed 'Data Broker' as a means to efficiently couple the components of our climate system model is described.

Farrara, John D.↗

From Oxygen Generation to Metals Production: In Situ Resource Utilization by Molten Oxide Electrolysis

For the exploration of other bodies in the solar system, electrochemical processing is arguably the most versatile technology for conversion of local resources into usable commodities: by electrolysis one can, in principle, produce (1) breathable oxygen, (2) silicon for the fabrication of solar cells, (3) various reactive metals for use as electrodes in advanced storage batteries, and (4) structural metals such as steel and aluminum. Even so, to date there has been no sustained effort to develop such processes, in part due to the inadequacy of the database. The objective here is to identify chemistries capable of sustaining molten oxide electrolysis in the cited applications and to examine the behavior of laboratory-scale cells designed to generate oxygen and to produce metal. The basic research includes the study of the underlying high-temperature physical chemistry of oxide melts representative of lunar regolith and of Martian soil. To move beyond empirical approaches to process development, the thermodynamic and transport properties of oxide melts are being studied to help set the limits of composition and temperature for the processing trials conducted in laboratory-scale electrolysis cells. The goal of this investigation is to deliver a working prototype cell that can use lunar regolith and Martian soil to produce breathable oxygen along with metal by-product. Additionally, the process can be generalized to permit adaptation to accommodate different feedstock chemistries, such as those that will be encountered on other bodies in the solar system. The expected results of this research include: (1) the identification of appropriate electrolyte chemistries; (2) the selection of candidate anode and cathode materials compatible with electrolytes named above; and (3) performance data from a laboratory-scale cell producing oxygen and metal. On the strength of these results it should be possible to assess the technical viability of molten oxide electrolysis for in situ resource utilization on the Moon and Mars. In parallel, there may be commercial applications here on earth, such as new green technologies for metals extraction and for treatment of hazardous waste, e.g., fixing heavy metals.

Khetpal, Deepak↗

Performance Assessment of OVERFLOW on Distributed Computing Environment

The aerodynamic computer code, OVERFLOW, with a multi-zone overset grid feature, has been parallelized to enhance its performance on distributed and shared memory paradigms. Practical application benchmarks have been set to assess the efficiency of code's parallelism on high-performance architectures. The code's performance has also been experimented with in the context of the distributed computing paradigm on distant computer resources using the Information Power Grid (IPG) toolkit, Globus. Two parallel versions of the code, namely OVERFLOW-MPI and -MLP, have developed around the natural coarse grained parallelism inherent in a multi-zonal domain decomposition paradigm. The algorithm invokes a strategy that forms a number of groups, each consisting of a zone, a cluster of zones and/or a partition of a large zone. Each group can be thought of as a process with one or multithreads assigned to it and that all groups run in parallel. The -MPI version of the code uses explicit message-passing based on the standard MPI library for sending and receiving interzonal boundary data across processors. The -MLP version employs no message-passing paradigm; the boundary data is transferred through the shared memory. The -MPI code is suited for both distributed and shared memory architectures, while the -MLP code can only be used on shared memory platforms. The IPG applications are implemented by the -MPI code using the Globus toolkit. While a computational task is distributed across multiple computer resources, the parallelism can be explored on each resource alone. Performance studies are achieved with some practical aerodynamic problems with complex geometries, consisting of 2.5 up to 33 million grid points and a large number of zonal blocks. The computations were executed primarily on SGI Origin 2000 multiprocessors and on the Cray T3E. OVERFLOW's IPG applications are carried out on NASA homogeneous metacomputing machines located at three sites, Ames, Langley and Glenn. Plans for the future will exploit the distributed parallel computing capability on various homogeneous and heterogeneous resources and large scale benchmarks. Alternative IPG toolkits will be used along with sophisticated zonal grouping strategies to minimize the communication time across the computer resources.

Djomehri, M. Jahed↗

Results of the Second U.S. Manned Suborbital Space Flight, July 21, 1961

This document presents the results of the second United States manned suborbital space flight. The data and flight description presented form a continuation of the information provided at an open conference held under the auspices of the National Aeronautics and Space Administration, in cooperation with the National Institutes of Health and the National Academy of Sciences, at the U.S. Department of State Auditorium on June 6, 1961. The papers presented herein generally parallel the presentations of the first report and were prepared by the personnel of the NASA Manned Spacecraft Center in collaboration with personnel from other government agencies, participating industry, and universities. The second successful manned suborbital space flight on July 21, 1961, in which Astronaut Virgil I. Grissom was the pilot was another step in the progressive research, development, and training program leading to the study of man's capabilities in a space environment during manned orbital flight. Data and operational experiences gained from this flight were in agreement with and supplemented the knowledge obtained from the first suborbital flight of May 5, 1961, piloted by Astronaut Alan B. Shepard, Jr. The two recent manned suborbital flights, coupled with the unmanned research and development flights, have provided valuable engineering nd scientific data on which the program can progress. The successful active participation of the pilots, in much the same way as in the development and testing of high performance aircraft, has. greatly increased our confidence in giving man a significant role in future space flight activities. It is the purpose of this report to continue the practice of providing data to the scientific community interested in activities of this nature. Brief descriptions are presented of the Project Mercury spacecraft and flight plan. Papers are provided which parallel the presentations of data published for the first suborbital space flight. Additional information is given relating to the operational aspects of the medical support activities for the two manned suborbital space flights.

Source record↗

Parametric analysis of RF communications and tracking systems for manned space stations

System performance, system interface compatibility, and system management and operations are analyzed for external and internal communications provided by the modular space station. Mathematical models are utilized to evaluate performances of the various communication links between the MSS and the tracking and data relay satellites, the shuttle, and the ground stations. Communication control design requirements are determined from overall MSS control design and checkout requirements are developed in a related systems study. Parallel design efforts consider equipment configurations for the external communication assembly baseband and the internal communication assembly to accommodate signal transfer and control. The recommended baseline design considers also interfaces between equipment groups and the critical functional and signal characteristics of these interfaces. The data processing assembly has direct communication circuits to the external communication assembly and the control interface is provided via a digital data bus and a remote acquisition and control unit.

Source record↗

Battery Reinitialization of the Photovoltaic Module of the International Space Station

The photovoltaic (PV) module on the International Space Station (ISS) has been operating since November 2000 and supporting electric power demands of the ISS and its crew of three. The PV module contains photovoltaic arrays that convert solar energy to electrical power and an integrated equipment assembly (IEA) that houses electrical hardware and batteries for electric power regulation and storage. Each PV module contains two independent power channels for fault tolerance. Each power channel contains three batteries in parallel to meet its performance requirements and for fault tolerance. Each battery consists of 76 Ni-Hydrogen (Ni-H2) cells in series. These 76 cells are contained in two orbital replaceable units (ORU) that are connected in series. On-orbit data are monitored and trended to ensure that all hardware is operating normally. Review of on-orbit data showed that while five batteries are operating very well, one is showing signs of mismatched ORUs. The cell pressure in the two ORUs differs by an amount that exceeds the recommended range. The reason for this abnormal behavior may be that the two ORUs have different use history. An assessment was performed and it was determined that capacity of this battery would be limited by the lower pressure ORU. Steps are being taken to reduce this pressure differential before battery capacity drops to the point of affecting its ability to meet performance requirements. As a first step, a battery reinitialization procedure was developed to reduce this pressure differential. The procedure was successfully carried out on-orbit and the pressure differential was reduced to the recommended range. This paper describes the battery performance and the consequences of mismatched ORUs that make a battery. The paper also describes the reinitialization procedure, how it was performed on orbit, and battery performance after the reinitialization. On-orbit data monitoring and trending is an ongoing activity and it will continue as ISS assembly progresses.

Hajela, Gyan↗

The Goes-R Geostationary Lightning Mapper (GLM): Algorithm and Instrument Status

The Geostationary Operational Environmental Satellite (GOES-R) is the next series to follow the existing GOES system currently operating over the Western Hemisphere. Superior spacecraft and instrument technology will support expanded detection of environmental phenomena, resulting in more timely and accurate forecasts and warnings. Advancements over current GOES capabilities include a new capability for total lightning detection (cloud and cloud-to-ground flashes) from the Geostationary Lightning Mapper (GLM), and improved capability for the Advanced Baseline Imager (ABI). The Geostationary Lighting Mapper (GLM) will map total lightning activity (in-cloud and cloud-to-ground lighting flashes) continuously day and night with near-uniform spatial resolution of 8 km with a product refresh rate of less than 20 sec over the Americas and adjacent oceanic regions. This will aid in forecasting severe storms and tornado activity, and convective weather impacts on aviation safety and efficiency. In parallel with the instrument development (a prototype and 4 flight models), a GOES-R Risk Reduction Team and Algorithm Working Group Lightning Applications Team have begun to develop the Level 2 algorithms, cal/val performance monitoring tools, and new applications. Proxy total lightning data from the NASA Lightning Imaging Sensor on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional test beds are being used to develop the pre-launch algorithms and applications, and also improve our knowledge of thunderstorm initiation and evolution. A joint field campaign with Brazilian researchers in 2010-2011 will produce concurrent observations from a VHF lightning mapping array, Meteosat multi-band imagery, Tropical Rainfall Measuring Mission (TRMM) Lightning Imaging Sensor (LIS) overpasses, and related ground and in-situ lightning and meteorological measurements in the vicinity of Sao Paulo. These data will provide a new comprehensive proxy data set for algorithm and application development.

Goodman, Steven J.↗

A Concept for Run-Time Support of the Chapel Language

A document presents a concept for run-time implementation of other concepts embodied in the Chapel programming language. (Now undergoing development, Chapel is intended to become a standard language for parallel computing that would surpass older such languages in both computational performance in the efficiency with which pre-existing code can be reused and new code written.) The aforementioned other concepts are those of distributions, domains, allocations, and access, as defined in a separate document called "A Semantic Framework for Domains and Distributions in Chapel" and linked to a language specification defined in another separate document called "Chapel Specification 0.3." The concept presented in the instant report is recognition that a data domain that was invented for Chapel offers a novel approach to distributing and processing data in a massively parallel environment. The concept is offered as a starting point for development of working descriptions of functions and data structures that would be necessary to implement interfaces to a compiler for transforming the aforementioned other concepts from their representations in Chapel source code to their run-time implementations.

James, Mark↗

Special purpose parallel computer architecture for real-time control and simulation in robotic applications

This is a real-time robotic controller and simulator which is a MIMD-SIMD parallel architecture for interfacing with an external host computer and providing a high degree of parallelism in computations for robotic control and simulation. It includes a host processor for receiving instructions from the external host computer and for transmitting answers to the external host computer. There are a plurality of SIMD microprocessors, each SIMD processor being a SIMD parallel processor capable of exploiting fine grain parallelism and further being able to operate asynchronously to form a MIMD architecture. Each SIMD processor comprises a SIMD architecture capable of performing two matrix-vector operations in parallel while fully exploiting parallelism in each operation. There is a system bus connecting the host processor to the plurality of SIMD microprocessors and a common clock providing a continuous sequence of clock pulses. There is also a ring structure interconnecting the plurality of SIMD microprocessors and connected to the clock for providing the clock pulses to the SIMD microprocessors and for providing a path for the flow of data and instructions between the SIMD microprocessors. The host processor includes logic for controlling the RRCS by interpreting instructions sent by the external host computer, decomposing the instructions into a series of computations to be performed by the SIMD microprocessors, using the system bus to distribute associated data among the SIMD microprocessors, and initiating activity of the SIMD microprocessors to perform the computations on the data by procedure call.

Fijany, Amir↗

Impact Testing of Inconel 718 for Material Impact Model Development

One of the difficulties with developing and verifying accurate impact models is that parameters such as high strain-rate material properties, failure modes, static properties, and impact test measurements are often obtained from a variety of different sources using different materials, with little control over consistency among the different sources. In addition, there is often a lack of quantitative measurements in impact tests to which the models can be compared. To alleviate some of these problems, a project is underway to develop a consistent set of material property, impact test data, and failure analysis for a variety of aircraft materials that can be used to develop improved impact failure and deformation models. This project is jointly funded by the NASA Glenn Research Center and the Federal Aviation Administration (FAA) William J. Hughes Technical Center. Unique features of this set of data are that all material property and impact test data are obtained using traceable material, the test methods and procedures are extensively documented, and all of the raw data is available. Four parallel efforts are currently underway. The Ohio State University conducts both measurement of material deformation and failure response over a wide range of strain rates and temperatures and failure analysis of material property specimens and impact test articles. The George Mason University conducts the development of improved numerical modeling techniques for deformation and failure. Glenn conducts the impact testing of flat panels and substructures. This report describes impact testing performed on Inconel 718 sheet and plate samples of different thicknesses with different types of projectiles, one a regular cylinder and one with a more complex geometry incorporating features representative of a jet engine fan blade. Data from this testing will be used in validating material models developed under this program. The material tests and the material models developed in this program will be published in separate reports.

J Michael Pereira↗

Applications Performance Under MPL and MPI on NAS IBM SP2

On July 5, 1994, an IBM Scalable POWER parallel System (IBM SP2) with 64 nodes, was installed at the Numerical Aerodynamic Simulation (NAS) Facility Each node of NAS IBM SP2 is a "wide node" consisting of a RISC 6000/590 workstation module with a clock of 66.5 MHz which can perform four floating point operations per clock with a peak performance of 266 Mflop/s. By the end of 1994, 64 nodes of IBM SP2 will be upgraded to 160 nodes with a peak performance of 42.5 Gflop/s. An overview of the IBM SP2 hardware is presented. The basic understanding of architectural details of RS 6000/590 will help application scientists the porting, optimizing, and tuning of codes from other machines such as the CRAY C90 and the Paragon to the NAS SP2. Optimization techniques such as quad-word loading, effective utilization of two floating point units, and data cache optimization of RS 6000/590 is illustrated, with examples giving performance gains at each optimization step. The conversion of codes using Intel's message passing library NX to codes using native Message Passing Library (MPL) and the Message Passing Interface (NMI) library available on the IBM SP2 is illustrated. In particular, we will present the performance of Fast Fourier Transform (FFT) kernel from NAS Parallel Benchmarks (NPB) under MPL and MPI. We have also optimized some of Fortran BLAS 2 and BLAS 3 routines, e.g., the optimized Fortran DAXPY runs at 175 Mflop/s and optimized Fortran DGEMM runs at 230 Mflop/s per node. The performance of the NPB (Class B) on the IBM SP2 is compared with the CRAY C90, Intel Paragon, TMC CM-5E, and the CRAY T3D.

Saini, Subhash↗

Solid propellant rocket motor internal ballistics performance variation analysis, phase 3

Results of research aimed at improving the predictability of off nominal internal ballistics performance of solid propellant rocket motors (SRMs) including thrust imbalance between two SRMs firing in parallel are reported. The potential effects of nozzle throat erosion on internal ballistic performance were studied and a propellant burning rate low postulated. The propellant burning rate model when coupled with the grain deformation model permits an excellent match between theoretical results and test data for the Titan IIIC, TU455.02, and the first Space Shuttle SRM (DM-1). Analysis of star grain deformation using an experimental model and a finite element model shows the star grain deformation effects for the Space Shuttle to be small in comparison to those of the circular perforated grain. An alternative technique was developed for predicting thrust imbalance without recourse to the Monte Carlo computer program. A scaling relationship used to relate theoretical results to test results may be applied to the alternative technique of predicting thrust imbalance or to the Monte Carlo evaluation. Extended investigation into the effect of strain rate on propellant burning rate leads to the conclusion that the thermoelastic effect is generally negligible for both steadily increasing pressure loads and oscillatory loads.

Sforzini, R. H.↗