Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing (computers)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

Automated Euler and Navier-Stokes Database Generation for a Glide-Back Booster

The past two decades have seen a sustained increase in the use of high fidelity Computational Fluid Dynamics (CFD) in basic research, aircraft design, and the analysis of post-design issues. As the fidelity of a CFD method increases, the number of cases that can be readily and affordably computed greatly diminishes. However, computer speeds now exceed 2 GHz, hundreds of processors are currently available and more affordable, and advances in parallel CFD algorithms scale more readily with large numbers of processors. All of these factors make it feasible to compute thousands of high fidelity cases. However, there still remains the overwhelming task of monitoring the solution process. This paper presents an approach to automate the CFD solution process. A new software tool, AeroDB, is used to compute thousands of Euler and Navier-Stokes solutions for a 2nd generation glide-back booster in one week. The solution process exploits a common job-submission grid environment, the NASA Information Power Grid (IPG), using 13 computers located at 4 different geographical sites. Process automation and web-based access to a MySql database greatly reduces the user workload, removing much of the tedium and tendency for user input errors. The AeroDB framework is shown. The user submits/deletes jobs, monitors AeroDB's progress, and retrieves data and plots via a web portal. Once a job is in the database, a job launcher uses an IPG resource broker to decide which computers are best suited to run the job. Job/code requirements, the number of CPUs free on a remote system, and queue lengths are some of the parameters the broker takes into account. The Globus software provides secure services for user authentication, remote shell execution, and secure file transfers over an open network. AeroDB automatically decides when a job is completed. Currently, the Cart3D unstructured flow solver is used for the Euler equations, and the Overflow structured overset flow solver is used for the Navier-Stokes equations. Other codes can be readily included into the AeroDB framework.

Chaderjian, Neal M.↗

Evolutionary Computational Methods for Identifying Emergent Behavior in Autonomous Systems

A technique based on Evolutionary Computational Methods (ECMs) was developed that allows for the automated optimization of complex computationally modeled systems, such as autonomous systems. The primary technology, which enables the ECM to find optimal solutions in complex search spaces, derives from evolutionary algorithms such as the genetic algorithm and differential evolution. These methods are based on biological processes, particularly genetics, and define an iterative process that evolves parameter sets into an optimum. Evolutionary computation is a method that operates on a population of existing computational-based engineering models (or simulators) and competes them using biologically inspired genetic operators on large parallel cluster computers. The result is the ability to automatically find design optimizations and trades, and thereby greatly amplify the role of the system engineer.

Terrile, Richard J.↗

Improving NASA's Multiscale Modeling Framework for Tropical Cyclone Climate Study

One of the current challenges in tropical cyclone (TC) research is how to improve our understanding of TC interannual variability and the impact of climate change on TCs. Recent advances in global modeling, visualization, and supercomputing technologies at NASA show potential for such studies. In this article, the authors discuss recent scalability improvement to the multiscale modeling framework (MMF) that makes it feasible to perform long-term TC-resolving simulations. The MMF consists of the finite-volume general circulation model (fvGCM), supplemented by a copy of the Goddard cumulus ensemble model (GCE) at each of the fvGCM grid points, giving 13,104 GCE copies. The original fvGCM implementation has a 1D data decomposition; the revised MMF implementation retains the 1D decomposition for most of the code, but uses a 2D decomposition for the massive copies of GCEs. Because the vast majority of computation time in the MMF is spent computing the GCEs, this approach can achieve excellent speedup without incurring the cost of modifying the entire code. Intelligent process mapping allows differing numbers of processes to be assigned to each domain for load balancing. The revised parallel implementation shows highly promising scalability, obtaining a nearly 80-fold speedup by increasing the number of cores from 30 to 3,335.

tropical cyclone interannual variability↗

Oculometric Analysis of Saccadic Compensation for Visual Motion Processing Impairment due to Alcohol and Sleep Disruption

The Visuomotor Control Laboratory at Ames Research Center has developed a 5-minute ocular tracking test that computes 21 largely independent metrics of visuomotor performance, reflecting neural signal processing along a number of distinct pathways through cortex, brainstem, and cerebellum. Human sensorimotor performance is resilient to the challenges and stressors of many operational environments, in part, because overall performance is achieved through multiple parallel systems. Our multidimensional oculometrics allow us to examine impacts on these sub-components separately. To illustrate this, we contrasted the effects of two mild neural stressors, acute sleep-deprivation and low-dose alcohol. We have previously shown that, in both cases, oculometric analysis is a highly sensitive indicator of impairment. Here we quantified not only the observed impact on the performance of one sub-system, smooth pursuit, which uses high-level cortical processing of visual motion to track a moving object, but also the observed (partial) compensation by an evolutionarily older mid-brain and brainstem subsystem, saccades, which generates jumps in eye position to catch up with the target when smooth pursuit is inadequate. Specifically, we examined the dose-response (effect size vs. dose size) of the ground lost (pursuit deficit) and the ground recouped (saccadic compensation) across three separate studies – acute low-dose alcohol administration (16 subjects), acute sleep loss (12 subjects), and acute sleep loss with caffeine intervention (9 subjects). We computed the dose-response slopes using linear regression. The figure below shows that, in the case of acute sleep deprivation, the resulting slopes for ground lost and ground recouped (mean ± SE across subjects) were significantly different (paired t-test, t(11) = 5.17, p < 0.001), indicating poor saccadic compensation. However, when sleep loss was coupled with caffeine ingestion, ground lost was decreased and ground recouped increased such that the slopes were no longer different (t(8) = -0.05, p = 0.965). With alcohol, the two slopes were large albeit not significantly different (t(15) = 0.96, p = 0.351), indicating significant pursuit impairment but effective saccadic compensation. Our findings show that sleep deprivation and alcohol affect oculomotor performance differently. Low-dose alcohol effects appear predominantly cortical, with effective brainstem compensation. Sleep loss and circadian disruption however appears to affect both cortical and brainstem pathways with caffeine providing an effective countermeasure to both effects. Beyond the mere detection of impairment, our oculometric assessment allows us to characterize the nature of the deficit, to provide insight into the neural substrate, and to assess the effectiveness of countermeasures.

pursuit↗

An interactive parallel programming environment applied in atmospheric science

This article introduces an interactive parallel programming environment (IPPE) that simplifies the generation and execution of parallel programs. One of the tasks of the environment is to generate message-passing parallel programs for homogeneous and heterogeneous computing platforms. The parallel programs are represented by using visual objects. This is accomplished with the help of a graphical programming editor that is implemented in Java and enables portability to a wide variety of computer platforms. In contrast to other graphical programming systems, reusable parts of the programs can be stored in a program library to support rapid prototyping. In addition, runtime performance data on different computing platforms is collected in a database. A selection process determines dynamically the software and the hardware platform to be used to solve the problem in minimal wall-clock time. The environment is currently being tested on a Grand Challenge problem, the NASA four-dimensional data assimilation system.

Interactive Display Devices↗

Multiprogramming and the performance of parallel programs

A programming methodology is introduced that utilizes computational synchronization and avoids tight control flow synchronization in parallel programs. In this methodology, each phase of the computation is assigned a status that can be ready, blocked, or completed, and tasks in each computational phase are self-scheduled to ensure computational progress by the available executing processes. Results indicate that this methodology avoids the catastrophic performance losses resulting from the swapping of processes in multiprogrammed multiprocessors.

Benten, Muhammad S.↗

C++ Resource Intelligent Compilation for GPU Enabled Applications

We are nearing the limits of Moore's Law with current computing technology. As industries push for more performance from smaller systems, alternate methods of computation such as Graphics Processing Units (GPUs) should be considered. Many of these systems utilize the Compute Unified Device Architecture (CUDA) to give programmers access to individual compute elements of the GPU for general purpose computing tasks. Direct access to the GPU's parallel multi-core architecture enables highly efficient computation and can drastically reduce the time required for complex algorithms or data analysis. Of course not all systems have a CUDA-enabled device to leverage, and so applications must consider optional support for users with these devices. Resource Intelligent Compilation (RIC) addresses this situation by enabling GPU-based acceleration of existing applications without affecting users without GPUs. Resource Intelligent Compilation (RIC) creates C/C++ modules that can be compiled to create a standard CPU version or GPU accelerated version of a program, depending on hardware availability. This is accomplished through a toolbox of programming strategies based on features of the CUDA API. Using this toolbox, existing applications can be modified with ease to support GPU acceleration, and new applications can be generated with just a few simple modifications. All of this culminates in an accelerated application for users with the appropriate hardware, with no performance impact to standard systems. This memorandum presents all the important features involved in supporting and implementing RIC and an example of using RIC to accelerate an existing mathematical model, without removing support for standard users. Through this memorandum, NASA engineers can acquire a set of guidelines to follow for RIC-compliant development, seamlessly accelerating C/C++ applications.

GPU↗

Mission-Maps For Outbound Cislunar Transfer Trajectories

This study quantifies the robustness and sensitivity of an outbound cislunar trajectory for a lunar lander in the form of mission-maps, or topological maps that allows either a computer program or mission designer to intuitively optimize the placement of critical outbound correction burns from the derived sensitivity data. The non-linear multi-body dynamics are applied to generate an outbound cislunar reference profile used by a linear covariance analysis (LinCov) tool to compute the expected Δv and trajectory dispersions due to the initial state uncertainty, sensor errors, maneuver execution errors, and disturbance accelerations along the outbound cislunar profile. The rapid performance analysis capabilities of LinCov are complimented with parallel processing techniques to evaluate hundreds and thousands of different translational burn locations, placements, and targeting constraints to identify the combination that minimizes the total Δv usage (nominal plus 3σ Δv) and trajectory dispersions at lunar orbit insertion. This study utilizes a generalized reference targeting algorithm to quickly assess the integrated closed-loop GN&C system performance due to different targeting configurations and constraints. The resulting mission maps provide an intuitive insight to ascertain each trajectory correction maneuver’s (TCM) sensitivity to different burn times along an outbound cislunar trajectory and quickly identify desirable engineering tradeoffs when performing analysis on the number and placement of these burns that nominally zero. Multiple mission maps are generated for a variety of different performance parameters that allow engineers to visually identify optimal solutions for trajectory correction maneuver placements, the number of correction burns, and the targeting constraints for each burn.

GN&C↗

Stability Analysis of the Flow Over a Swept Forward-Facing Step Using PIV Base Flows

Step excrescences are a type of surface imperfection encountered on swept wings of commercial aircraft, commonly because accessibility requirements prevent creating the wing’s surface from a single panel. If the step height is too large, super-critical, the flow can undergo an early transition to turbulence, which can be highly detrimental for the performance of the wing. In the case of a forward-facing step on a swept wing in a low-disturbance environment, stationary crossflow vortices can develop a significant amplitude upstream of the step and hence dominate the structure of the boundary-layer flow over the step. The main goal of the present investigation is to illuminate the path to transition supported by the fascinatingly complex flow field in the direct downstream vicinity of a step with a super critical height. The high-resolution, stereographic Particle Image Velocimetry (PIV) measurement dataset presently available for this flow field provides a complete description of the laminar flow for the execution of BiGlobal stability analysis in a plane parallel to the step. Although the notorious sensitivity of stability results to the description of the base flow demands a very careful uncertainty analysis of those results, it is argued that this very fact can be leveraged to produce new insight into the supported perturbation dynamics. In performing the analysis, several unsteady mode families are discovered that display the explosive perturbation expected for early transition to be induced. In considering domain widths equal to an integer-multiple of the incident crossflow-vortex wavelength and analyzing an extent of 5 crossflow-vortex wavelengths parallel to the step, it is found that the stability results converge while increasing the domain width. It is demonstrated, moreover, that the results for the wider domains can be approximated by appropriately averaging the results on neighboring single-crossflow-vortex-wavelength domains covering the same region. Besides being useful for computational purposes, this observed property suggests interpreting the instability mechanism as a distorted primary mechanism rather than a “proper” secondary mechanism. This follows in the context of the secondary in-stability analysis of three-dimensional boundary layers, because the secondary mechanism is usually characterized by being localized in a pocket of strong shear, while the distorted primary mechanism typically has an infinite support in the direction parallel to the step. Even though the growth rates are found to be sensitive to the interrogation-window size inherent to the PIV post-processing procedure, the spatial structure of the eigenfunctions is found to be relatively insensitive. Lastly, the spatial structure of the eigenfunctions corresponding to all velocity components are matched with the shape functions determined by computing the Spectral Proper Orthogonal Decomposition (SPOD) of a time-resolved measurement of the perturbation content.

forward-facing step↗

A survey on the design of multiprocessing systems for artificial intelligence applications

Some issues in designing computers for artificial intelligence (AI) processing are discussed. These issues are divided into three levels: the representation level, the control level, and the processor level. The representation level deals with the knowledge and methods used to solve the problem and the means to represent it. The control level is concerned with the detection of dependencies and parallelism in the algorithmic and program representations of the problem, and with the synchronization and sheduling of concurrent tasks. The processor level addresses the hardware and architectural components needed to evaluate the algorithmic and program representations. Solutions for the problems of each level are illustrated by a number of representative systems. Design decisions in existing projects on AI computers are classed into top-down, bottom-up, and middle-out approaches.

Wah, Benjamin W.↗

An open, parallel I/O computer as the platform for high-performance, high-capacity mass storage systems

APTEC Computer Systems is a Portland, Oregon based manufacturer of I/O computers. APTEC's work in the context of high density storage media is on programs requiring real-time data capture with low latency processing and storage requirements. An example of APTEC's work in this area is the Loral/Space Telescope-Data Archival and Distribution System. This is an existing Loral AeroSys designed system, which utilizes an APTEC I/O computer. The key attributes of a system architecture that is suitable for this environment are as follows: (1) data acquisition alternatives; (2) a wide range of supported mass storage devices; (3) data processing options; (4) data availability through standard network connections; and (5) an overall system architecture (hardware and software designed for high bandwidth and low latency). APTEC's approach is outlined in this document.

Abineri, Adrian↗

Parallel eigenanalysis of finite element models in a completely connected architecture

A parallel algorithm is presented for the solution of the generalized eigenproblem in linear elastic finite element analysis, (K)(phi) = (M)(phi)(omega), where (K) and (M) are of order N, and (omega) is order of q. The concurrent solution of the eigenproblem is based on the multifrontal/modified subspace method and is achieved in a completely connected parallel architecture in which each processor is allowed to communicate with all other processors. The algorithm was successfully implemented on a tightly coupled multiple-instruction multiple-data parallel processing machine, Cray X-MP. A finite element model is divided into m domains each of which is assumed to process n elements. Each domain is then assigned to a processor or to a logical processor (task) if the number of domains exceeds the number of physical processors. The macrotasking library routines are used in mapping each domain to a user task. Computational speed-up and efficiency are used to determine the effectiveness of the algorithm. The effect of the number of domains, the number of degrees-of-freedom located along the global fronts and the dimension of the subspace on the performance of the algorithm are investigated. A parallel finite element dynamic analysis program, p-feda, is documented and the performance of its subroutines in parallel environment is analyzed.

Akl, F. A.↗

Simulating futures in extended common LISP

Stack-groups comprise the mechanism underlying implementation of multiprocessing in Extended Common LISP, i.e., running multiple quasi-simultaneous processes within a single LISP address space. On the other hand, the future construct of MULTILISP, an extension of the LISP dialect scheme, deals with parallel execution. The source of concurrency that future exploits is the overlap between computation of a value and use of the value. Described is a simulation of the future construct by an interpreter utilizing stack-group extensions to common LISP.

Nachtsheim, Philip R.↗

Aeroelastic problems in turbomachines

A review of the field of turbomachinery aeroelasticity is presented. Developments over the past decade are emphasized, and an assessment of possible future directions of research is offered. The paper reviews the areas of unsteady cascade flows, structural modeling, and flutter prediction methods. Representative results for unsteady flow calculations and flutter boundary predictions in subsonic, transonic, and supersonic flows are discussed, including recent calculations based on the methods of computational fluid mechanics. Results from current attempts to correlate experimental data with theoretical predictions are discussed briefly. It is recommended that future research include investigations of novel approaches to flutter calculations that can take full advantage of parallel processing supercomputers. The feasibility of using mistuning and aeroelastic tailoring as passive flutter suppression techniques should also be pursued.

Bendiksen, Oddvar O.↗

Hyperswitch communication network

The Hyperswitch Communication Network (HCN) is a large scale parallel computer prototype being developed at JPL. Commercial versions of the HCN computer are planned. The HCN computer being designed is a message passing multiple instruction multiple data (MIMD) computer, and offers many advantages in price-performance ratio, reliability and availability, and manufacturing over traditional uniprocessors and bus based multiprocessors. The design of the HCN operating system is a uniquely flexible environment that combines both parallel processing and distributed processing. This programming paradigm can achieve a balance among the following competing factors: performance in processing and communications, user friendliness, and fault tolerance. The prototype is being designed to accommodate a maximum of 64 state of the art microprocessors. The HCN is classified as a distributed supercomputer. The HCN system is described, and the performance/cost analysis and other competing factors within the system design are reviewed.

Peterson, J.↗

Benchmarking Memory Performance with the Data Cube Operator

Data movement across a computer memory hierarchy and across computational grids is known to be a limiting factor for applications processing large data sets. We use the Data Cube Operator on an Arithmetic Data Set, called ADC, to benchmark capabilities of computers and of computational grids to handle large distributed data sets. We present a prototype implementation of a parallel algorithm for computation of the operatol: The algorithm follows a known approach for computing views from the smallest parent. The ADC stresses all levels of grid memory and storage by producing some of 2d views of an Arithmetic Data Set of d-tuples described by a small number of integers. We control data intensity of the ADC by selecting the tuple parameters, the sizes of the views, and the number of realized views. Benchmarking results of memory performance of a number of computer architectures and of a small computational grid are presented.

Frumkin, Michael A.↗

Using the GeoFEST Faulted Region Simulation System

GeoFEST (the Geophysical Finite Element Simulation Tool) simulates stress evolution, fault slip and plastic/elastic processes in realistic materials, and so is suitable for earthquake cycle studies in regions such as Southern California. Many new capabilities and means of access for GeoFEST are now supported. New abilities include MPI-based cluster parallel computing using automatic PYRAMID/Parmetis-based mesh partitioning, automatic mesh generation for layered media with rectangular faults, and results visualization that is integrated with remote sensing data. The parallel GeoFEST application has been successfully run on over a half-dozen computers, including Intel Xeon clusters, Itanium II and Altix machines, and the Apple G5 cluster. It is not separately optimized for different machines, but relies on good domain partitioning for load-balance and low communication, and careful writing of the parallel diagonally preconditioned conjugate gradient solver to keep communication overhead low. Demonstrated thousand-step solutions for over a million finite elements on 64 processors require under three hours, and scaling tests show high efficiency when using more than (order of) 4000 elements per processor. The source code and documentation for GeoFEST is available at no cost from Open Channel Foundation. In addition GeoFEST may be used through a browser-based portal environment available to approved users. That environment includes semi-automated geometry creation and mesh generation tools, GeoFEST, and RIVA-based visualization tools that include the ability to generate a flyover animation showing deformations and topography. Work is in progress to support simulation of a region with several faults using 16 million elements, using a strain energy metric to adapt the mesh to faithfully represent the solution in a region of widely varying strain.

Geophyical Finite Element Simulation Tool (GeoFEST↗

Magnetic Reconfiguration in Explosive Solar Activity

A fundamental property of the Sun's corona i s that it is violently dynamic. The most spectacular and most energetic manifestations of this activity are the giant disruptions that give rise to coronal mass ejections (CME) and eruptive flares. These major events are of critical importance, because they drive the most destructive forms of space weather at Earth and in the solar system, and they provide a unique opportunity to study, in revealing detail, the interaction of magnetic field and matter, in particular, magnetohydrodynamic instability and nonequilibrium -- processes that are at the heart of laboratory and astrophysical plasma physics. Recent observations by a number of NASA space missions have given us new insights into the physical mechanisms that underlie coronal explosions. Furthermore, massively-parallel computation have now allowed us to calculate fully three-dimensional models for solar activity. In this talk I will present some of the latest observations of the Sun, including those from the just-launched Hinode and STEREO mission, and discuss recent advances in the theory and modeling of explosive solar activity.

Antiochos, Spiro K.↗