Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Models of Wake-Vortex Spreading Mechanisms and Their Estimated Uncertainties

One of the primary constraints on the capacity of the nation's air transportation system is the landing capacity at its busiest airports. Many airports with nearly-simultaneous operations on closely-spaced parallel runways (i.e., as close as 750 ft (246m)) suffer a severe decrease in runway acceptance rate when weather conditions do not allow full utilization. The objective of a research program at NASA Ames Research Center is to develop the technologies needed for traffic management in the airport environment so that operations now allowed on closely-spaced parallel runways under Visual Meteorological Conditions can also be carried out under Instrument Meteorological Conditions. As part of this overall research objective, the study reported here has developed improved models for the various aerodynamic mechanisms that spread and transport wake vortices. The purpose of the study is to continue the development of relationships that increase the accuracy of estimates for the along-trail separation distances available before the vortex wake of a leading aircraft intrudes into the airspace of a following aircraft. Details of the models used and their uncertainties are presented in the appendices to the paper. Suggestions are made as to the theoretical and experimental research needed to increase the accuracy of and confidence level in the models presented and instrumentation required or more precise estimates of the motion and spread of vortex wakes. The improved wake models indicate that, if the following aircraft is upwind of the leading aircraft, the vortex wakes of the leading aircraft will not intrude into the airspace of the following aircraft for about 7s (based on pessimistic assumptions) for most atmospheric conditions. The wake-spreading models also indicate that longer time intervals before wake intrusion are available when atmospheric turbulence levels are mild or moderate. However, if the estimates for those time intervals are to be reliable, further study is necessary to develop the instrumentation and procedures needed to accurately define when the more benign atmospheric conditions exist.

Rossow, Vernon J.↗

Shared Memory Parallelization of an Implicit ADI-type CFD Code

A parallelization study designed for ADI-type algorithms is presented using the OpenMP specification for shared-memory multiprocessor programming. Details of optimizations specifically addressed to cache-based computer architectures are described and performance measurements for the single and multiprocessor implementation are summarized. The paper demonstrates that optimization of memory access on a cache-based computer architecture controls the performance of the computational algorithm. A hybrid MPI/OpenMP approach is proposed for clusters of shared memory machines to further enhance the parallel performance. The method is applied to develop a new LES/DNS code, named LESTool. A preliminary DNS calculation of a fully developed channel flow at a Reynolds number of 180, Re(sub tau) = 180, has shown good agreement with existing data.

Hauser, Th.↗

Extending substructure based iterative solvers to multiple load and repeated analyses

Direct solvers currently dominate commercial finite element structural software, but do not scale well in the fine granularity regime targeted by emerging parallel processors. Substructure based iterative solvers--often called also domain decomposition algorithms--lend themselves better to parallel processing, but must overcome several obstacles before earning their place in general purpose structural analysis programs. One such obstacle is the solution of systems with many or repeated right hand sides. Such systems arise, for example, in multiple load static analyses and in implicit linear dynamics computations. Direct solvers are well-suited for these problems because after the system matrix has been factored, the multiple or repeated solutions can be obtained through relatively inexpensive forward and backward substitutions. On the other hand, iterative solvers in general are ill-suited for these problems because they often must restart from scratch for every different right hand side. In this paper, we present a methodology for extending the range of applications of domain decomposition methods to problems with multiple or repeated right hand sides. Basically, we formulate the overall problem as a series of minimization problems over K-orthogonal and supplementary subspaces, and tailor the preconditioned conjugate gradient algorithm to solve them efficiently. The resulting solution method is scalable, whereas direct factorization schemes and forward and backward substitution algorithms are not. We illustrate the proposed methodology with the solution of static and dynamic structural problems, and highlight its potential to outperform forward and backward substitutions on parallel computers. As an example, we show that for a linear structural dynamics problem with 11640 degrees of freedom, every time-step beyond time-step 15 is solved in a single iteration and consumes 1.0 second on a 32 processor iPSC-860 system; for the same problem and the same parallel processor, a pair of forward/backward substitutions at each step consumes 15.0 seconds.

Farhat, Charbel↗

Integrating Cache Performance Modeling and Tuning Support in Parallelization Tools

With the resurgence of distributed shared memory (DSM) systems based on cache-coherent Non Uniform Memory Access (ccNUMA) architectures and increasing disparity between memory and processors speeds, data locality overheads are becoming the greatest bottlenecks in the way of realizing potential high performance of these systems. While parallelization tools and compilers facilitate the users in porting their sequential applications to a DSM system, a lot of time and effort is needed to tune the memory performance of these applications to achieve reasonable speedup. In this paper, we show that integrating cache performance modeling and tuning support within a parallelization environment can alleviate this problem. The Cache Performance Modeling and Prediction Tool (CPMP), employs trace-driven simulation techniques without the overhead of generating and managing detailed address traces. CPMP predicts the cache performance impact of source code level "what-if" modifications in a program to assist a user in the tuning process. CPMP is built on top of a customized version of the Computer Aided Parallelization Tools (CAPTools) environment. Finally, we demonstrate how CPMP can be applied to tune a real Computational Fluid Dynamics (CFD) application.

Waheed, Abdul↗

Combinatorial transcription factor binding encodes cis -regulatory wiring of mouse forebrain GABAergic neurogenesis

Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.

59 BASIC BIOLOGICAL SCIENCES↗

Implementation of Helioseismic Data Reduction and Diagnostic Techniques on Massively Parallel Architectures

Under the direction of Dr. Rhodes, and the technical supervision of Dr. Korzennik, the data assimilation of high spatial resolution solar dopplergrams has been carried out throughout the program on the Intel Delta Touchstone supercomputer. With the help of a research assistant, partially supported by this grant, and under the supervision of Dr. Korzennik, code development was carried out at SAO, using various available resources. To ensure cross-platform portability, PVM was selected as the message passing library. A parallel implementation of power spectra computation for helioseismology data reduction, using PVM was successfully completed. It was successfully ported to SMP architectures (i.e. SUN), and to some MPP architectures (i.e. the CM5). Due to limitation of the implementation of PVM on the Cray T3D, the port to that architecture was not completed at the time.

Korzennik, Sylvain↗

Beyond the Renderer: Software Architecture for Parallel Graphics and Visualization

As numerous implementations have demonstrated, software-based parallel rendering is an effective way to obtain the needed computational power for a variety of challenging applications in computer graphics and scientific visualization. To fully realize their potential, however, parallel renderers need to be integrated into a complete environment for generating, manipulating, and delivering visual data. We examine the structure and components of such an environment, including the programming and user interfaces, rendering engines, and image delivery systems. We consider some of the constraints imposed by real-world applications and discuss the problems and issues involved in bringing parallel rendering out of the lab and into production.

Crockett, Thomas W.↗

A multicomputer and real time Ada environment

A multicomputer is defined as a set of tightly-coupled yet autonomous computers capable of synchronizing and communicating in parallel but also of operating independently. The architectural concepts and requirements for executing the Ada programs in a multicomputer system are discussed. Synchronization, commmunication, and protection of shared data between the Ada program entities are addressed. Decomposition or partitioning of the Ada program in a multicomputer system is also studied. Finally, a multicomputer and real time Ada environment is described using FLEX/32 multicomputer system.

Naeini, Ray↗

Structural stability augmentation system design using BODEDIRECT: A quick and accurate approach

A methodology is presented for a modal suppression control law design using flight test data instead of mathematical models to obtain the required gain and phase information about the flexible airplane. This approach is referred to as BODEDIRECT. The purpose of the BODEDIRECT program is to provide a method of analyzing the modal phase relationships measured directly from the airplane. These measurements can be achieved with a frequency sweep at the control surface input while measuring the outputs of interest. The measured Bode-models can be used directly for analysis in the frequency domain, and for control law design. Besides providing a more accurate representation for the system inputs and outputs of interest, this method is quick and relatively inexpensive. To date, the BODEDIRECT program has been tested and verified for computational integrity. Its capabilities include calculation of series, parallel and loop closure connections between Bode-model representations. System PSD, together with gain and phase margins of stability may be calculated for successive loop closures of multi-input/multi-output systems. Current plans include extensive flight testing to obtain a Bode-model representation of a commercial aircraft for design of a structural stability augmentation system.

Goslin, T. J.↗

Hypercluster parallel processing library user's manual

This User's Manual describes the Hypercluster Parallel Processing Library, composed of FORTRAN-callable subroutines which enable a FORTRAN programmer to manipulate and transfer information throughout the Hypercluster at NASA Lewis Research Center. Each subroutine and its parameters are described in detail. A simple heat flow application using Laplace's equation is included to demonstrate the use of some of the library's subroutines. The manual can be used initially as an introduction to the parallel features provided by the library. Thereafter it can be used as a reference when programming an application.

Quealy, Angela↗

Shared versus distributed memory multiprocessors

The question of whether multiprocessors should have shared or distributed memory has attracted a great deal of attention. Some researchers argue strongly for building distributed memory machines, while others argue just as strongly for programming shared memory multiprocessors. A great deal of research is underway on both types of parallel systems. Special emphasis is placed on systems with a very large number of processors for computation intensive tasks and considers research and implementation trends. It appears that the two types of systems will likely converge to a common form for large scale multiprocessors.

Jordan, Harry F.↗

Temperature dependence of diffusivities, preliminary definition phase

During the six months definition phase of the instrument development program, research personnel at the Center for Microgravity and Materials Research of the University of Alabama in Huntsville (UAH) were to furnish all of the necessary labor, services, materials, and facilities necessary to provide science requirement definition, initiate hardware development activities, requirements and timetable for integration and experimental accommodation of the GAS payload into the Shuttle cargo bay and an updated ground-based research flight program proposal consistent with the NRA selection letter. These activities were to be accomplished in parallel and consistent with the necessary research and development work toward the accomplishment of the overall objectives of the selected proposal.

Rosenberger, Franz↗

Cloud Spatial Structure and 3D Radiative Transfer

Cloud radiative properties are sensitive to drop size and other parameters of cloud micro-structure, but also to cloud shape,spacing, and other parameters of cloud macro-structure, including internal fractal structure. New information on cloud structure is being derived from a variety of cloud radars, and ongoing field programs such as DoE/ARM. These programs are improving the measurement and modelling of physical and radiative properties of clouds. A parallel effort is underway to improve cloud remote sensing, especially from the new suite of EOS-AM1 instruments which will provide higher spectral, spatial resolution, and/or angular resolution. Key parameters for improving pixel-scale retrievals are cloud thickness and photon mean-free-path, which together determine the scale of "radiative smoothing" of cloud fluxes and radiances. This scale has been observed as a change in the spatial spectrum of Landsat cloud radiances, and was also recently found with the Goddard micropulse lidar, by searching for returns from directions nonparallel to the incident beam. "Offbeam" Lidar returns are now being used to estimate the cloud "radiative Green's function", G,which depends on cloud thickness and may be used to retrieve that important quantity. G is also being applied to improving simple IPA estimates of cloud radiative properties. This and other measurements of 3D transfer in clouds, coupled with Monte Carlo and other 3D transfer methods, are beginning to provide a better understanding of the dependence of radiation on cloud inhomogeneity, and to suggest new retrieval and parameterization algorithms which take account of cloud inhomogeneity.

Cahalan, Robert F.↗

The 'Biologically-Inspired Computing' Column

The field of Biology changed dramatically in 1953, with the determination by Francis Crick and James Dewey Watson of the double helix structure of DNA. This discovery changed Biology for ever, allowing the sequencing of the human genome, and the emergence of a "new Biology" focused on DNA, genes, proteins, data, and search. Computational Biology and Bioinformatics heavily rely on computing to facilitate research into life and development. Simultaneously, an understanding of the biology of living organisms indicates a parallel with computing systems: molecules in living cells interact, grow, and transform according to the "program" dictated by DNA. Moreover, paradigms of Computing are emerging based on modelling and developing computer-based systems exploiting ideas that are observed in nature. This includes building into computer systems self-management and self-governance mechanisms that are inspired by the human body's autonomic nervous system, modelling evolutionary systems analogous to colonies of ants or other insects, and developing highly-efficient and highly-complex distributed systems from large numbers of (often quite simple) largely homogeneous components to reflect the behaviour of flocks of birds, swarms of bees, herds of animals, or schools of fish. This new field of "Biologically-Inspired Computing", often known in other incarnations by other names, such as: Autonomic Computing, Pervasive Computing, Organic Computing, Biomimetics, and Artificial Life, amongst others, is poised at the intersection of Computer Science, Engineering, Mathematics, and the Life Sciences. Successes have been reported in the fields of drug discovery, data communications, computer animation, control and command, exploration systems for space, undersea, and harsh environments, to name but a few, and augur much promise for future progress.

Hinchey, Mike↗

Dividing the Concentrator Target From the Genesis Mission

The Genesis spacecraft, launched in 2001, traveled to a Lagrangian point between the Earth and Sun to collect particles from the solar wind and return them to Earth. However, during the return of the spacecraft in 2004, the parachute failed to open during descent, and the Genesis spacecraft crashed into the Utah desert. Many of the solar wind collectors were broken into smaller pieces, and the field team rapidly collected the capsule and collector pieces for later assessment. On each of the next few days, the team discovered that various collectors had survived intact, including three of four concentrator targets. Within a month, the team had imaged more than 10,000 fragments and packed them for transport to the Astromaterials Acquisition and Curation Office within the ARES Directorate at JSC. Currently, the Genesis samples are curated along with the other extraterrestrial sample collections within ARES. Although they were broken and dirty, the Genesis solar wind collectors still offered the science community the opportunity to better understand our Sun and the solar system as a whole. One of the more highly prized concentrator collectors survived the crash almost completely intact. The Genesis Concentrator was designed to concentrate the solar wind by a factor of at least 20 so that solar oxygen and nitrogen isotopes could be measured. One of these materials was the Diamond-on-Silicon (DoS) concentrator target. Unfortunately, the DoS concentrator broke on impact. Nevertheless, the scientific value of the DoS concentrator target was high. The Genesis Allocation Committee received a request for approximately 1 cm(sup 2) of the DoS specimen taken near the focal point of the concentrator for the analysis of solar wind nitrogen isotopes. The largest fragment, Genesis sample 60000, was designated for this allocation and needed to be precisely cut. The requirement was to subdivide the designated sample in a manner that prevented contamination of the sample and minimized the risk of losing or breaking the precious requested sample fragment. The Genesis curator determined that the use of laser scribing techniques to "cut" a precise line and subsequently cleave the sample (in a controlled break of the sample along that line) was the best method for accomplishing the sample subdivision. However, there were risks, including excess heating of the sample, that could cause some of the implanted solar wind to be lost via thermal diffusion. Accidentally breaking the sample during the handling and cleaving process was an additional risk. Early in fiscal year 2013, to address this delicate, complicated task, the ARES Directorate assembled its top scientists to develop a cutting plan that would ensure success when applied to the actual concentrator target wafer; i.e., to produce an approximately 1 cm(sup 2) piece from the requested area of the wafer. The team, subsequently referred to as the JSC Genesis Tiger Team, spent months researching and testing parameters and techniques related to scribing, cleaving, transporting, handling, and holding (i.e., mounting) the specimen. The investigation required considerable "thinking outside the box," and many, many trials using nonflight wafer analogs. After all preliminary testing, the following method was adopted as the final cutting plan. It was used in two final end-to-end practice runs before being used on the actual flight target wafer. The wafer was oriented on the laser cutting stage with the 100 and 010 directions of the wafer parallel to the corresponding X and Y directions of the cutting stage. The laser was programed to scribe 31 lines of the appropriate length along the Y stage direction. The programed scribe lines were separated by 5 micron in the X direction. The laser parameters were set as follows: (1) The laser power was 0.5 watts; (2) each line consisted of 50 passes, with the Z position being advanced 5 micron per pass; and (3) 30 s would elapse before the next line was scribed to allow for wafer cool down from any possible heating via the laser. The ablated material that "stuck" in the "scribe-cut" was removed from the "cut" using an ultrasonic micro-tool. After all the ablated silicon was removed from the wafer, the wafer was repositioned in exactly the same orientation on the laser stage. The laser was focused using the bottom of the wafer channel, and the 31-line scribing pattern described above was reprogrammed using the Z position of the groove bottom as the starting Z value instead of the top wafer surface, which was used previously. Upon completion of the second set of scribes, the ultrasonic micro-tool was again used to clean out the cut. The wafer was remounted on the stage in exactly the same orientation as before. The laser was again focused on the bottom of the groove. This time, however, the laser was.programed to scribe only one line down the exact center of the channel. The final scribe line consisted of 100 passes with a Z advance of 5 micron per pass and with the laser power set at 0.5 watts. As mentioned above, the final cutting plan was practiced in two end-to-end trials using non-flight, triangular-shaped silicon wafers similar in size and orientation to the actual DOS 60000 target sample. The actual scribing of the triangular-shaped wafers required scribing two lines and cleaving (i.e. scribe-cleave, then scribe-cleave) to obtain the piece requested for allocation. Early in December 2012, after many months of experiments and practicing and perfecting the techniques and procedures, the team successfully subdivided the Genesis DoS 60000 target sample, one of the most scientifically important samples from the Genesis mission (figure 2). On December 17, 2012, the allocated piece of concentrator target sample was delivered to the requesting principal investigator.The cutting plan developed for the subdivision of this sample will be used as the model for subdividing future requested Genesis flight wafers (appropriately modified for different wafer types).

Lauer, H. V., Jr.↗

SafeDNN: Understanding and Verifying Neural Networks

The SafeDNN project at NASA Ames explores analysis techniques and tools to ensure that systems that use Deep Neural Networks (DNN) are safe, robust and interpretable. Research directions we are pursuing include: symbolic execution for DNN analysis, label-guided clustering to automatically identify input regions that are robust, parallel and compositional approaches to improve formal SMT-based verification, property inference and automated program repair for DNNs, adversarial training and detection, probabilistic reasoning for DNNs. In this talk I will highlight some of the research advances from SafeDNN, that were already published.

Corina Pasareanu↗

Massively Parallel Dantzig-Wolfe Decomposition Applied to Traffic Flow Scheduling

Optimal scheduling of air traffic over the entire National Airspace System is a computationally difficult task. To speed computation, Dantzig-Wolfe decomposition is applied to a known linear integer programming approach for assigning delays to flights. The optimization model is proven to have the block-angular structure necessary for Dantzig-Wolfe decomposition. The subproblems for this decomposition are solved in parallel via independent computation threads. Experimental evidence suggests that as the number of subproblems/threads increases (and their respective sizes decrease), the solution quality, convergence, and runtime improve. A demonstration of this is provided by using one flight per subproblem, which is the finest possible decomposition. This results in thousands of subproblems and associated computation threads. This massively parallel approach is compared to one with few threads and to standard (non-decomposed) approaches in terms of solution quality and runtime. Since this method generally provides a non-integral (relaxed) solution to the original optimization problem, two heuristics are developed to generate an integral solution. Dantzig-Wolfe followed by these heuristics can provide a near-optimal (sometimes optimal) solution to the original problem hundreds of times faster than standard (non-decomposed) approaches. In addition, when massive decomposition is employed, the solution is shown to be more likely integral, which obviates the need for an integerization step. These results indicate that nationwide, real-time, high fidelity, optimal traffic flow scheduling is achievable for (at least) 3 hour planning horizons.

Rios, Joseph Lucio↗