Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

An analysis of I/O efficient order-statistic-based techniques for noise power estimation in the HRMS sky survey's operational system

Noise power estimation in the High-Resolution Microwave Survey (HRMS) sky survey element is considered as an example of a constant false alarm rate (CFAR) signal detection problem. Order-statistic-based noise power estimators for CFAR detection are considered in terms of required estimator accuracy and estimator dynamic range. By limiting the dynamic range of the value to be estimated, the performance of an order-statistic estimator can be achieved by simpler techniques requiring only a single pass of the data. Simple threshold-and-count techniques are examined, and it is shown how several parallel threshold-and-count estimation devices can be used to expand the dynamic range to meet HRMS system requirements with minimal hardware complexity. An input/output (I/O) efficient limited-precision order-statistic estimator with wide but limited dynamic range is also examined.

Zimmerman, G. A.↗

Lead-Free Experiment in a Space Environment

This Technical Memorandum addresses the Lead-Free Technology Experiment in Space Environment that flew as part of the seventh Materials International Space Station Experiment outside the International Space Station for approximately 18 months. Its intent was to provide data on the performance of lead-free electronics in an actual space environment. Its postflight condition is compared to the preflight condition as well as to the condition of an identical package operating in parallel in the laboratory. Some tin whisker growth was seen on a flight board but the whiskers were few and short. There were no solder joint failures, no tin pest formation, and no significant intermetallic compound formation or growth on either the flight or ground units.

Blanche, J. F.↗

Fast Simulation of the NICER Instrument

The NICER mission uses a complicated physical system to collect information from objects that are, by x-ray timing science standards, rather faint. To get the most out of the data we will need a rigorous understanding of all instrumental effects. We are in the process of constructing a very fast, high fidelity simulator that will help us to assess instrument performance, support simulation-based data reduction, and improve our estimates of measurement error. We will combine and extend existing optics, detector, and electronics simulations. We will employ the Compute Unied Device Architecture (CUDA2) to parallelize these calculations. The price of suitable CUDA-compatible multi-gigaflop cores is about $0.20/core, so this approach will be very cost-effective.

gEDA↗

VLSI system for synthetic aperture radar /SAR/ processing

The paper discusses an SAR problem based on actual requirements set forth by NASA for a spaceborne application. The requirements for high resolution and high quality necessitate a data sampling rate of 7.5 MHz. For each data value 1,025 4-bit complex multiply + add operations are needed, which is equivalent to 7.7 GHz complex multiply + add operation rate. Since this rate is much too high for general purpose systems, a special-purpose device was sought. This paper discusses two architectures based on parallel operation of 1,025 identical cells, each of which is capable of performing arithmetic, storage, and several control operations. The operation rate in each device is only 7.5 MHz, which is quite manageable, especially with the help of a substantial degree of pipelining. A computational-mathematical analysis is used as a primary tool for evaluating the design and some of its tradeoffs. Two different approaches are discussed and compared; both are based on having 1,025 identical cells working in parallel, but differ in their dual approaches to the flow of data.

Cohen, D.↗

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry↗

Mechanical characterization of PMR-15 graphite/polyimide bolted joints

Data are presented for the static and viscoelastic performance of Celion 6000/PMR-15 graphite/polyimide composite belted joints over the 21 to 315 C temperature range. Two laminate configurations were obtained by sectioning a panel into subpanels oriented parallel and perpendicular to the laminate principal direction. Three replicates of each specimen geometry were tested at each of the test temperatures for the two laminate configurations, and failure load was defined as the maximum load attained as indicated by the load/deflection curve. Effects of strain-rate on strength and failure mode behavior were assessed at room temperature for strain rates of 0.002, 0.01, 0.10, and 1.00/s, and creep behavior was examined for the case of bearing deformations at temperatures of 21 and 177 C. Results are presented, including: (1) two distinct mechanisms were responsible for bearing failure and related to laminate configuration, (2) the + or - 45 deg plies carried the shearing load in composite bolted joints, and (3) the material exhibited strain-rate sensitivity which varied with laminate configuration.

Wilson, D. W.↗

Ice crystal and ice nucleus measurements in cap clouds

Ice nucleation in cap clouds over a mountain in Wyoming was examined with airborne instrumentation. Crosswind and wind parallel passes were made through the clouds, with data being taken on the ice crystal concentrations and sizes. A total of 141 penetrations of 26 separate days in temperatures ranging from -7 to -24 C were performed. Subsequent measurements were also made 100 km away from the mountain. The ice crystal concentrations measured showed good correlation with the ice nucleus content in winter time, midcontinental air masses in Wyoming. Further studies are recommended to determine if the variations in the ice nucleus population are the cause of the variability if ice crystal content.

Vali, G.↗

A parallel simulated annealing algorithm for standard cell placement on a hypercube computer

A parallel version of a simulated annealing algorithm is presented which is targeted to run on a hypercube computer. A strategy for mapping the cells in a two dimensional area of a chip onto processors in an n-dimensional hypercube is proposed such that both small and large distance moves can be applied. Two types of moves are allowed: cell exchanges and cell displacements. The computation of the cost function in parallel among all the processors in the hypercube is described along with a distributed data structure that needs to be stored in the hypercube to support parallel cost evaluation. A novel tree broadcasting strategy is used extensively in the algorithm for updating cell locations in the parallel environment. Studies on the performance of the algorithm on example industrial circuits show that it is faster and gives better final placement results than the uniprocessor simulated annealing algorithms. An improved uniprocessor algorithm is proposed which is based on the improved results obtained from parallelization of the simulated annealing algorithm.

Jones, Mark Howard↗

Redundant disk arrays: Reliable, parallel secondary storage

During the past decade, advances in processor and memory technology have given rise to increases in computational performance that far outstrip increases in the performance of secondary storage technology. Coupled with emerging small-disk technology, disk arrays provide the cost, volume, and capacity of current disk subsystems, by leveraging parallelism, many times their performance. Unfortunately, arrays of small disks may have much higher failure rates than the single large disks they replace. Redundant arrays of inexpensive disks (RAID) use simple redundancy schemes to provide high data reliability. The data encoding, performance, and reliability of redundant disk arrays are investigated. Organizing redundant data into a disk array is treated as a coding problem. Among alternatives examined, codes as simple as parity are shown to effectively correct single, self-identifying disk failures.

Gibson, Garth Alan↗

Neural networks for calibration tomography

Artificial neural networks are suitable for performing pattern-to-pattern calibrations. These calibrations are potentially useful for facilities operations in aeronautics, the control of optical alignment, and the like. Computed tomography is compared with neural net calibration tomography for estimating density from its x-ray transform. X-ray transforms are measured, for example, in diffuse-illumination, holographic interferometry of fluids. Computed tomography and neural net calibration tomography are shown to have comparable performance for a 10 degree viewing cone and 29 interferograms within that cone. The system of tomography discussed is proposed as a relevant test of neural networks and other parallel processors intended for using flow visualization data.

Decker, Arthur↗

A MIMD implementation of a parallel Euler solver for unstructured grids

A mesh-vertex finite volume scheme for solving the Euler equations on triangular unstructured meshes is implemented on a MIMD (multiple instruction/multiple data stream) parallel computer. Three partitioning strategies for distributing the work load onto the processors are discussed. Issues pertaining to the communication costs are also addressed. We find that the spectral bisection strategy yields the best performance. The performance of this unstructured computation on the Intel iPSC/860 compares very favorably with that on a one-processor CRAY Y-MP/1 and an earlier implementation on the Connection Machine.

Venkatakrishnan, V.↗

Preliminary Examination of the Interstellar Collector of Stardust

The findings of the Stardust spacecraft mission returned to earth in January 2006 are discussed. The spacecraft returned two unprecedented and independent extraterrestrial samples: the first sample of a comet and the first samples of contemporary interstellar dust. An important lesson from the cometary Preliminary Examination (PE) was that the Stardust cometary samples in aerogel presented a technical challenge. Captured particles often separate into multiple fragments, intimately mix with aerogel and are typically buried hundreds of microns to millimeters deep in the aerogel collectors. The interstellar dust samples are likely much more challenging since they are expected to be orders of magnitudes smaller in mass, and their fluence is two orders of magnitude smaller than that of the cometary particles. The goal of the Stardust Interstellar Preliminary Examination (ISPE) is to answer several broad questions, including: which features in the interstellar collector aerogel were generated by hypervelocity impact and how much morphological and trajectory information may be gained?; how well resolved are the trajectories of probable interstellar particles from those of interplanetary origin?; and, by comparison to impacts by known particle dimensions in laboratory experiments, what was the mass distribution of the impacting particles? To answer these questions, and others, non-destructive, sequential, non-invasive analyses of interstellar dust candidates extracted from the Stardust interstellar tray will be performed. The total duration of the ISPE will be three years and will differ from the Stardust cometary PE in that data acquisition for the initial characterization stage will be prolonged and will continue simultaneously and parallel with data publications and release of the first samples for further investigation.

Westphal, A. J.↗

A Novel Method for Quantifying Helmeted Field of View of a Space Suit - And What it Means for Constellation

Field of view has always been a design feature paramount to helmet design, and in particular spacesuit design, where the helmet must provide an adequate field of view for a large range of activities, environments, and body positions. Historically, suited field of view has been evaluated either qualitatively in parallel with design or quantitatively using various test methods and protocols. As such, oftentimes legacy suit field of view information is either ambiguous for lack of supporting data or contradictory to other field of view tests performed with different subjects and test methods. This paper serves to document a new field of view testing method that is more reliable and repeatable than its predecessors. It borrows heavily from standard field of vision tests such as the Goldmann kinetic perimetry test, but is designed specifically for evaluating field of view of a spacesuit helmet. In this test, three suits utilizing three different helmet designs were tested for field of view. Not only do these tests provide more reliable field of view data for legacy and prototype helmet designs, they also provide insight into how helmet design impacts field of view and what this means for the Constellation Project spacesuit helmet, which must meet stringent field of view requirements that are more generous to the crewmember than legacy designs.

McFarland, Shane M.↗

Grid Task Execution

IPG Execution Service is a framework that reliably executes complex jobs on a computational grid, and is part of the IPG service architecture designed to support location-independent computing. The new grid service enables users to describe the platform on which they need a job to run, which allows the service to locate the desired platform, configure it for the required application, and execute the job. After a job is submitted, users can monitor it through periodic notifications, or through queries. Each job consists of a set of tasks that performs actions such as executing applications and managing data. Each task is executed based on a starting condition that is an expression of the states of other tasks. This formulation allows tasks to be executed in parallel, and also allows a user to specify tasks to execute when other tasks succeed, fail, or are canceled. The two core components of the Execution Service are the Task Database, which stores tasks that have been submitted for execution, and the Task Manager, which executes tasks in the proper order, based on the user-specified starting conditions, and avoids overloading local and remote resources while executing tasks.

Hu, Chaumin↗

A Novel Method for Quantifying Helmeted Field of View of a Spacesuit - And What It Means for Constellation

Field of view has always been a design feature paramount to helmet design, and in particular spacesuit design, where the helmet must provide an adequate field of view for a large range of activities, environments, and body positions. Historically, suited field of view has been evaluated either qualitatively in parallel with design or quantitatively using various test methods and protocols. As such, oftentimes legacy suit field of view information is either ambiguous for lack of supporting data or contradictory to other field of view tests performed with different subjects and test methods. This paper serves to document a new field of view testing method that is more reliable and repeatable than its predecessors. It borrows heavily from standard ophthalmologic field of vision tests such as the Goldmann kinetic perimetry test, but is designed specifically for evaluating field of view of a spacesuit helmet. In this test, four suits utilizing three different helmet designs were tested for field of view. Not only do these tests provide more reliable field of view data for legacy and prototype helmet designs, they also provide insight into how helmet design impacts field of view and what this means for the Constellation Project spacesuit helmet, which must meet stringent field of view requirements that are more generous to the crewmember than legacy designs.

McFarland, Shane M.↗

Implementing Atmospheric Infrared Sounder (AIRS) and Cross-Track Infrared Sounder (CrIS) Cloud-Clearing Algorithm into the NASA GEOS: Focus on the 2017 Atlantic Tropical Cyclone Season

Numerical Weather Prediction (NWP) centers assimilate cloud-free infrared (IR) radiances because the assimilation of all-sky IR radiances is not yet operationally achievable. The cloud-clearing procedure offers a simpler, but effective strategy that produces cloud-affected radiances suitable for assimilation in partially cloudy regions. Several studies conducted by this team have demonstrated that IR Cloud-Cleared Radiances (CCRs), if thinned more aggressively than clear-sky radiances, can improve analysis and forecasts, particularly in meteorologically active areas. However, CCRs are not used by operational centers due partly to the thought that the process of cloud-clearing may affect latency and introduce difficult-to-control external dependencies. This study presents the results of implementing an Atmospheric Infrared Sounder (AIRS) and Cross-Track Infrared Sounder (CrIS) cloud-clearing procedure into the NASA Goddard Earth Observing System (GEOS) to demonstrate the portability of the procedure. The AIRS and CrIS cloud-clearing algorithms have been deprived of external dependencies, made customizable to any specific model, and the computational efficiency has been improved via parallelization. The revised AIRS and CrIS cloud-clearing algorithms allow a customized choice of channel selection, the use of a user-specified model's fields as first guess, and can perform in real time. Data assimilation experiments with the hybrid 4DEnVar GEOS system were successfully performed for the 2017 tropical cyclones (TC) season with a focus on three major hurricanes (Harvey, Irma, and Maria). This study shows that assimilation of locally-generated CCRs have a positive impact on both global skill and TC representation, compared to the assimilation of AIRS and CrIS clear-sky radiances, and a comparable or slightly improved impact compared to assimilation of CCRs produced by external sources, such as NASA's Distributed Active Archive Centers and NOAA’s Comprehensive Large Array-data Stewardship System. The customization and computational efficiency of the revised procedure would enable its usability in a real-time forecast context.

Niama Boukachaba↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Crowley, Kay↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Saltz, Joel↗