Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Integrating Cache Performance Modeling and Tuning Support in Parallelization Tools

With the resurgence of distributed shared memory (DSM) systems based on cache-coherent Non Uniform Memory Access (ccNUMA) architectures and increasing disparity between memory and processors speeds, data locality overheads are becoming the greatest bottlenecks in the way of realizing potential high performance of these systems. While parallelization tools and compilers facilitate the users in porting their sequential applications to a DSM system, a lot of time and effort is needed to tune the memory performance of these applications to achieve reasonable speedup. In this paper, we show that integrating cache performance modeling and tuning support within a parallelization environment can alleviate this problem. The Cache Performance Modeling and Prediction Tool (CPMP), employs trace-driven simulation techniques without the overhead of generating and managing detailed address traces. CPMP predicts the cache performance impact of source code level "what-if" modifications in a program to assist a user in the tuning process. CPMP is built on top of a customized version of the Computer Aided Parallelization Tools (CAPTools) environment. Finally, we demonstrate how CPMP can be applied to tune a real Computational Fluid Dynamics (CFD) application.

Waheed, Abdul↗

Combinatorial transcription factor binding encodes cis -regulatory wiring of mouse forebrain GABAergic neurogenesis

Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.

59 BASIC BIOLOGICAL SCIENCES↗

Implementation of Helioseismic Data Reduction and Diagnostic Techniques on Massively Parallel Architectures

Under the direction of Dr. Rhodes, and the technical supervision of Dr. Korzennik, the data assimilation of high spatial resolution solar dopplergrams has been carried out throughout the program on the Intel Delta Touchstone supercomputer. With the help of a research assistant, partially supported by this grant, and under the supervision of Dr. Korzennik, code development was carried out at SAO, using various available resources. To ensure cross-platform portability, PVM was selected as the message passing library. A parallel implementation of power spectra computation for helioseismology data reduction, using PVM was successfully completed. It was successfully ported to SMP architectures (i.e. SUN), and to some MPP architectures (i.e. the CM5). Due to limitation of the implementation of PVM on the Cray T3D, the port to that architecture was not completed at the time.

Korzennik, Sylvain↗

Beyond the Renderer: Software Architecture for Parallel Graphics and Visualization

As numerous implementations have demonstrated, software-based parallel rendering is an effective way to obtain the needed computational power for a variety of challenging applications in computer graphics and scientific visualization. To fully realize their potential, however, parallel renderers need to be integrated into a complete environment for generating, manipulating, and delivering visual data. We examine the structure and components of such an environment, including the programming and user interfaces, rendering engines, and image delivery systems. We consider some of the constraints imposed by real-world applications and discuss the problems and issues involved in bringing parallel rendering out of the lab and into production.

Crockett, Thomas W.↗

A multicomputer and real time Ada environment

A multicomputer is defined as a set of tightly-coupled yet autonomous computers capable of synchronizing and communicating in parallel but also of operating independently. The architectural concepts and requirements for executing the Ada programs in a multicomputer system are discussed. Synchronization, commmunication, and protection of shared data between the Ada program entities are addressed. Decomposition or partitioning of the Ada program in a multicomputer system is also studied. Finally, a multicomputer and real time Ada environment is described using FLEX/32 multicomputer system.

Naeini, Ray↗

Structural stability augmentation system design using BODEDIRECT: A quick and accurate approach

A methodology is presented for a modal suppression control law design using flight test data instead of mathematical models to obtain the required gain and phase information about the flexible airplane. This approach is referred to as BODEDIRECT. The purpose of the BODEDIRECT program is to provide a method of analyzing the modal phase relationships measured directly from the airplane. These measurements can be achieved with a frequency sweep at the control surface input while measuring the outputs of interest. The measured Bode-models can be used directly for analysis in the frequency domain, and for control law design. Besides providing a more accurate representation for the system inputs and outputs of interest, this method is quick and relatively inexpensive. To date, the BODEDIRECT program has been tested and verified for computational integrity. Its capabilities include calculation of series, parallel and loop closure connections between Bode-model representations. System PSD, together with gain and phase margins of stability may be calculated for successive loop closures of multi-input/multi-output systems. Current plans include extensive flight testing to obtain a Bode-model representation of a commercial aircraft for design of a structural stability augmentation system.

Goslin, T. J.↗

Hypercluster parallel processing library user's manual

This User's Manual describes the Hypercluster Parallel Processing Library, composed of FORTRAN-callable subroutines which enable a FORTRAN programmer to manipulate and transfer information throughout the Hypercluster at NASA Lewis Research Center. Each subroutine and its parameters are described in detail. A simple heat flow application using Laplace's equation is included to demonstrate the use of some of the library's subroutines. The manual can be used initially as an introduction to the parallel features provided by the library. Thereafter it can be used as a reference when programming an application.

Quealy, Angela↗

Shared versus distributed memory multiprocessors

The question of whether multiprocessors should have shared or distributed memory has attracted a great deal of attention. Some researchers argue strongly for building distributed memory machines, while others argue just as strongly for programming shared memory multiprocessors. A great deal of research is underway on both types of parallel systems. Special emphasis is placed on systems with a very large number of processors for computation intensive tasks and considers research and implementation trends. It appears that the two types of systems will likely converge to a common form for large scale multiprocessors.

Jordan, Harry F.↗

Temperature dependence of diffusivities, preliminary definition phase

During the six months definition phase of the instrument development program, research personnel at the Center for Microgravity and Materials Research of the University of Alabama in Huntsville (UAH) were to furnish all of the necessary labor, services, materials, and facilities necessary to provide science requirement definition, initiate hardware development activities, requirements and timetable for integration and experimental accommodation of the GAS payload into the Shuttle cargo bay and an updated ground-based research flight program proposal consistent with the NRA selection letter. These activities were to be accomplished in parallel and consistent with the necessary research and development work toward the accomplishment of the overall objectives of the selected proposal.

Rosenberger, Franz↗

Cloud Spatial Structure and 3D Radiative Transfer

Cloud radiative properties are sensitive to drop size and other parameters of cloud micro-structure, but also to cloud shape,spacing, and other parameters of cloud macro-structure, including internal fractal structure. New information on cloud structure is being derived from a variety of cloud radars, and ongoing field programs such as DoE/ARM. These programs are improving the measurement and modelling of physical and radiative properties of clouds. A parallel effort is underway to improve cloud remote sensing, especially from the new suite of EOS-AM1 instruments which will provide higher spectral, spatial resolution, and/or angular resolution. Key parameters for improving pixel-scale retrievals are cloud thickness and photon mean-free-path, which together determine the scale of "radiative smoothing" of cloud fluxes and radiances. This scale has been observed as a change in the spatial spectrum of Landsat cloud radiances, and was also recently found with the Goddard micropulse lidar, by searching for returns from directions nonparallel to the incident beam. "Offbeam" Lidar returns are now being used to estimate the cloud "radiative Green's function", G,which depends on cloud thickness and may be used to retrieve that important quantity. G is also being applied to improving simple IPA estimates of cloud radiative properties. This and other measurements of 3D transfer in clouds, coupled with Monte Carlo and other 3D transfer methods, are beginning to provide a better understanding of the dependence of radiation on cloud inhomogeneity, and to suggest new retrieval and parameterization algorithms which take account of cloud inhomogeneity.

Cahalan, Robert F.↗

The 'Biologically-Inspired Computing' Column

The field of Biology changed dramatically in 1953, with the determination by Francis Crick and James Dewey Watson of the double helix structure of DNA. This discovery changed Biology for ever, allowing the sequencing of the human genome, and the emergence of a "new Biology" focused on DNA, genes, proteins, data, and search. Computational Biology and Bioinformatics heavily rely on computing to facilitate research into life and development. Simultaneously, an understanding of the biology of living organisms indicates a parallel with computing systems: molecules in living cells interact, grow, and transform according to the "program" dictated by DNA. Moreover, paradigms of Computing are emerging based on modelling and developing computer-based systems exploiting ideas that are observed in nature. This includes building into computer systems self-management and self-governance mechanisms that are inspired by the human body's autonomic nervous system, modelling evolutionary systems analogous to colonies of ants or other insects, and developing highly-efficient and highly-complex distributed systems from large numbers of (often quite simple) largely homogeneous components to reflect the behaviour of flocks of birds, swarms of bees, herds of animals, or schools of fish. This new field of "Biologically-Inspired Computing", often known in other incarnations by other names, such as: Autonomic Computing, Pervasive Computing, Organic Computing, Biomimetics, and Artificial Life, amongst others, is poised at the intersection of Computer Science, Engineering, Mathematics, and the Life Sciences. Successes have been reported in the fields of drug discovery, data communications, computer animation, control and command, exploration systems for space, undersea, and harsh environments, to name but a few, and augur much promise for future progress.

Hinchey, Mike↗

Dividing the Concentrator Target From the Genesis Mission

The Genesis spacecraft, launched in 2001, traveled to a Lagrangian point between the Earth and Sun to collect particles from the solar wind and return them to Earth. However, during the return of the spacecraft in 2004, the parachute failed to open during descent, and the Genesis spacecraft crashed into the Utah desert. Many of the solar wind collectors were broken into smaller pieces, and the field team rapidly collected the capsule and collector pieces for later assessment. On each of the next few days, the team discovered that various collectors had survived intact, including three of four concentrator targets. Within a month, the team had imaged more than 10,000 fragments and packed them for transport to the Astromaterials Acquisition and Curation Office within the ARES Directorate at JSC. Currently, the Genesis samples are curated along with the other extraterrestrial sample collections within ARES. Although they were broken and dirty, the Genesis solar wind collectors still offered the science community the opportunity to better understand our Sun and the solar system as a whole. One of the more highly prized concentrator collectors survived the crash almost completely intact. The Genesis Concentrator was designed to concentrate the solar wind by a factor of at least 20 so that solar oxygen and nitrogen isotopes could be measured. One of these materials was the Diamond-on-Silicon (DoS) concentrator target. Unfortunately, the DoS concentrator broke on impact. Nevertheless, the scientific value of the DoS concentrator target was high. The Genesis Allocation Committee received a request for approximately 1 cm(sup 2) of the DoS specimen taken near the focal point of the concentrator for the analysis of solar wind nitrogen isotopes. The largest fragment, Genesis sample 60000, was designated for this allocation and needed to be precisely cut. The requirement was to subdivide the designated sample in a manner that prevented contamination of the sample and minimized the risk of losing or breaking the precious requested sample fragment. The Genesis curator determined that the use of laser scribing techniques to "cut" a precise line and subsequently cleave the sample (in a controlled break of the sample along that line) was the best method for accomplishing the sample subdivision. However, there were risks, including excess heating of the sample, that could cause some of the implanted solar wind to be lost via thermal diffusion. Accidentally breaking the sample during the handling and cleaving process was an additional risk. Early in fiscal year 2013, to address this delicate, complicated task, the ARES Directorate assembled its top scientists to develop a cutting plan that would ensure success when applied to the actual concentrator target wafer; i.e., to produce an approximately 1 cm(sup 2) piece from the requested area of the wafer. The team, subsequently referred to as the JSC Genesis Tiger Team, spent months researching and testing parameters and techniques related to scribing, cleaving, transporting, handling, and holding (i.e., mounting) the specimen. The investigation required considerable "thinking outside the box," and many, many trials using nonflight wafer analogs. After all preliminary testing, the following method was adopted as the final cutting plan. It was used in two final end-to-end practice runs before being used on the actual flight target wafer. The wafer was oriented on the laser cutting stage with the 100 and 010 directions of the wafer parallel to the corresponding X and Y directions of the cutting stage. The laser was programed to scribe 31 lines of the appropriate length along the Y stage direction. The programed scribe lines were separated by 5 micron in the X direction. The laser parameters were set as follows: (1) The laser power was 0.5 watts; (2) each line consisted of 50 passes, with the Z position being advanced 5 micron per pass; and (3) 30 s would elapse before the next line was scribed to allow for wafer cool down from any possible heating via the laser. The ablated material that "stuck" in the "scribe-cut" was removed from the "cut" using an ultrasonic micro-tool. After all the ablated silicon was removed from the wafer, the wafer was repositioned in exactly the same orientation on the laser stage. The laser was focused using the bottom of the wafer channel, and the 31-line scribing pattern described above was reprogrammed using the Z position of the groove bottom as the starting Z value instead of the top wafer surface, which was used previously. Upon completion of the second set of scribes, the ultrasonic micro-tool was again used to clean out the cut. The wafer was remounted on the stage in exactly the same orientation as before. The laser was again focused on the bottom of the groove. This time, however, the laser was.programed to scribe only one line down the exact center of the channel. The final scribe line consisted of 100 passes with a Z advance of 5 micron per pass and with the laser power set at 0.5 watts. As mentioned above, the final cutting plan was practiced in two end-to-end trials using non-flight, triangular-shaped silicon wafers similar in size and orientation to the actual DOS 60000 target sample. The actual scribing of the triangular-shaped wafers required scribing two lines and cleaving (i.e. scribe-cleave, then scribe-cleave) to obtain the piece requested for allocation. Early in December 2012, after many months of experiments and practicing and perfecting the techniques and procedures, the team successfully subdivided the Genesis DoS 60000 target sample, one of the most scientifically important samples from the Genesis mission (figure 2). On December 17, 2012, the allocated piece of concentrator target sample was delivered to the requesting principal investigator.The cutting plan developed for the subdivision of this sample will be used as the model for subdividing future requested Genesis flight wafers (appropriately modified for different wafer types).

Lauer, H. V., Jr.↗

SafeDNN: Understanding and Verifying Neural Networks

The SafeDNN project at NASA Ames explores analysis techniques and tools to ensure that systems that use Deep Neural Networks (DNN) are safe, robust and interpretable. Research directions we are pursuing include: symbolic execution for DNN analysis, label-guided clustering to automatically identify input regions that are robust, parallel and compositional approaches to improve formal SMT-based verification, property inference and automated program repair for DNNs, adversarial training and detection, probabilistic reasoning for DNNs. In this talk I will highlight some of the research advances from SafeDNN, that were already published.

Corina Pasareanu↗

Massively Parallel Dantzig-Wolfe Decomposition Applied to Traffic Flow Scheduling

Optimal scheduling of air traffic over the entire National Airspace System is a computationally difficult task. To speed computation, Dantzig-Wolfe decomposition is applied to a known linear integer programming approach for assigning delays to flights. The optimization model is proven to have the block-angular structure necessary for Dantzig-Wolfe decomposition. The subproblems for this decomposition are solved in parallel via independent computation threads. Experimental evidence suggests that as the number of subproblems/threads increases (and their respective sizes decrease), the solution quality, convergence, and runtime improve. A demonstration of this is provided by using one flight per subproblem, which is the finest possible decomposition. This results in thousands of subproblems and associated computation threads. This massively parallel approach is compared to one with few threads and to standard (non-decomposed) approaches in terms of solution quality and runtime. Since this method generally provides a non-integral (relaxed) solution to the original optimization problem, two heuristics are developed to generate an integral solution. Dantzig-Wolfe followed by these heuristics can provide a near-optimal (sometimes optimal) solution to the original problem hundreds of times faster than standard (non-decomposed) approaches. In addition, when massive decomposition is employed, the solution is shown to be more likely integral, which obviates the need for an integerization step. These results indicate that nationwide, real-time, high fidelity, optimal traffic flow scheduling is achievable for (at least) 3 hour planning horizons.

Rios, Joseph Lucio↗

Programming substructure computations for elliptic problems on the CHiP system

A number of studies have been conducted with the aim to apply parallel computation to problems associated with solving finite element equations arising in structural mechanics and fluid dynamics. These studies have provided many important results. The present investigation is concerned with a set of experiments designed to test two ideas, including configurability and substructuring. The considered algorithms and tests are intended for implementation on the Configurable, Highly Parallel (CHiP) family of architecture described by Snyder (1982). The ChiP computer is composed of homogeneous processing elements (PEs) placed at regular intervals in a lattice of programmable switches. Two examples of the role of configurability and substructuring for simple iterative algorithms are considered, giving attention to conjugate gradient iterations, and tridiagonal systems of equations.

Gannon, D.↗

Automated Generation of Message-Passing Programs: An Evaluation of CAPTools using NAS Benchmarks

Scientists at NASA Ames Research Center have been developing computational aeroscience applications on highly parallel architectures over the past ten years. During the same time period, a steady transition of hardware and system software also occurred, forcing us to expand great efforts into migrating and receding our applications. As applications and machine architectures continue to become increasingly complex, the cost and time required for this process will become prohibitive. Various attempts to exploit software tools to assist and automate the parallelization process have not produced favorable results. In this paper, we evaluate an interactive parallelization tool, CAPTools, for parallelizing serial versions of the NAB Parallel Benchmarks. Finally, we compare the performance of the resulting CAPTools generated code to the hand-coded benchmarks on the Origin 2000 and IBM SP2. Based on these results, a discussion on the feasibility of automated parallelization of aerospace applications is presented along with suggestions for future work.

Hribar, Michelle R.↗

Electromagnetic diffraction by plane reflection diffraction gratings

A plane wave theory was developed to study electromagnetic diffraction by plane reflection diffraction gratings of infinite extent. A computer program was written to calculate the energy distribution in the various orders of diffraction for the cases when the electric or magnetic field vectors are parallel to the grating grooves. Within the region of validity of this theory, results were in excellent agreement with those in the literature. Energy conservation checks were also made to determine the region of validity of the plane wave theory. The computer program was flexible enough to analyze any grating profile that could be described by a single value function f(x). Within the region of validity the program could be used with confidence. The computer program was used to investigate the polarization and blaze properties of the diffraction grating.

Bocker, R. P.↗