Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

CLIPS meets the connection machine: Or how to create a parallel production system

Production systems usually present unacceptable run-times when faced with applications requiring tens of thousands to millions of facts. Many efforts have focused on the use of parallelism as a way to increase overall system performance. While these efforts have increased pattern matching and rule evaluation rates, they have only indirectly dealt with the problems faced by fact burdened applications. We have implemented PPS, a version of CLIPS running on the Connection Machine, to directly address the problems faced by these applications. This paper will describe our system, discuss its implementation, and present results.

Geyer, Steve↗

Two-Dimensional Sequential and Concurrent Finite Element Analysis of Unstiffened and Stiffened Aluminum and Composite Panels with Hole

The results of a detailed investigation of the distribution of stresses in aluminum and composite panels subjected to uniform end shortening are presented. The focus problem is a rectangular panel with two longitudinal stiffeners, and an inner stiffener discontinuous at a central hole in the panel. The influence of the stiffeners on the stresses is evaluated through a two-dimensional global finite element analysis in the absence or presence of the hole. Contrary to the physical feel, it is found that the maximum stresses from the glocal analysis for both stiffened aluminum and composite panels are greater than the corresponding stresses for the unstiffened panels. The inner discontinuous stiffener causes a greater increase in stresses than the reduction provided by the two outer stiffeners. A detailed layer-by-layer study of stresses around the hole is also presented for both unstiffened and stiffened composite panels. A parallel equation solver is used for the global system of equations since the computational time is far less than that using a sequential scheme. A parallel Choleski method with up to 16 processors is used on Flex/32 Multicomputer at NASA Langley Research Center. The parallel computing results are summarized and include the computational times, speedups, bandwidths, and their inter-relationships for the panel problems. It is found that the computational time for the Choleski method decreases with a decrease in bandwidth, and better speedups result as the bandwidth increases.

Razzaq, Zia↗

The role of analysis error in the convergence of reanalysis production streams in MERRA-2

Due to production time constraints, most reanalyses are produced in multiple parallel streams instead of a single continuous one. These streams cover separate segments of the reanalysis time period with short overlaps to allow reconstruction of the official record. A fundamental assumption justifying this approach is that the streams will be assimilating the same observations during the periods where they overlap, and so will eventually converge to a similar atmospheric state, making discontinuities at stream junctions negligible. This assumption is revisited in this work by examining the impact of analysis error on the differences between MERRA-2 overlapping streams in three historical periods. Comparison results are shown in terms of standard deviations of stream differences as well as the spectral decomposition of the variance of their differences. Residual differences were found at the end of each year of overlap, with larger values observed in the earlier segments of the presatellite era. By drawing parallels with analysis error statistics estimated from the GMAO OSSE system, these differences are shown to reflect the varying constraint of data with the varying observing network, and to further carry the imprint of errors that the data assimilation process is not able to mitigate. As such, they are unlikely to be reduced by longer spinup periods. The ability of data assimilation to ensure continuity in the parallel streams is put into question when the observing system coverage is inadequate or simply when the data assimilation system as a whole is suboptimal.

Amal El Akkraoui↗

Compositional changes, time and density variations in the magnetosphere associated with Birkeland currents and particle acceleration

Birkeland currents, parallel electric fields and plasma instabilities often occur together in time and space and play an important role in large scale plasma motions between the ionosphere and magnetosphere and along the magnetic field in the magnetosphere. The results of the plasma movements are large density and composition variations in the ionosphere and magnetosphere. Observations from ISIS-2 at 1400 km altitude show large densities with heavy ions dominating in regions with upward Birkeland currents, and low densities and light ions in regions with downward currents. Observations from ISEE-1 in field aligned current regions at 10,000 to 15,000 km altitude show transverse heating of protons and oxygen ions to 250 eV. Because of the different mobility of the protons and oxygen ions the proton flow is important in the beginning of the events but later the outflow becomes almost pure oxygen. Similarly ISEE-1 observations of outgoing field ion beams at 10,000 to 15,000 km altitude show time variations in the H+/O+ ratios and a dominance of O+ later in the events.

Ungstrup, E.↗

Execution time support for scientific programs on distributed memory machines

Optimizations are considered that are required for efficient execution of code segments that consists of loops over distributed data structures. The PARTI (Parallel Automated Runtime Toolkit at ICASE) execution time primitives are designed to carry out these optimizations and can be used to implement a wide range of scientific algorithms on distributed memory machines. These primitives allow the user to control array mappings in a way that gives an appearance of shared memory. Computations can be based on a global index set. Primitives are used to carry out gather and scatter operations on distributed arrays. Communications patterns are derived at runtime, and the appropriate send and receive messages are automatically generated.

Berryman, Harry↗

Optimal mapping of neural-network learning on message-passing multicomputers

A minimization of learning-algorithm completion time is sought in the present optimal-mapping study of the learning process in multilayer feed-forward artificial neural networks (ANNs) for message-passing multicomputers. A novel approximation algorithm for mappings of this kind is derived from observations of the dominance of a parallel ANN algorithm over its communication time. Attention is given to both static and dynamic mapping schemes for systems with static and dynamic background workloads, as well as to experimental results obtained for simulated mappings on multicomputers with dynamic background workloads.

Chu, Lon-Chan↗

Application of the hypercube parallel processor to a large-scale moment method code

The applicability of a parallel computing architecture to the solution of a large-scale moment-method code is investigated. Specifically, the NEC (Numerical Electromagnetics Code) method-of-moments scattering program is implemented on a hypercube parallel processor. The accuracy and the increase in the speed of execution on this parallel architecture are demonstrated. The results show a very large reduction in execution time for large problems. The great potential of this parallel processor is shown for interactive solution of large NEC problems as well as other moment-method techniques such as the finite-element method.

Manshadi, Farzin↗

Power-MOSFET Voltage Regulator

Ninety-six parallel MOSFET devices with two-stage feedback circuit form a high-current dc voltage regulator that also acts as fully-on solid-state switch when fuel-cell out-put falls below regulated voltage. Ripple voltage is less than 20 mV, transient recovery time is less than 50 ms. Parallel MOSFET's act as high-current dc regulator and switch. Regulator can be used wherever large direct currents must be controlled. Can be applied to inverters, industrial furnaces photovoltaic solar generators, dc motors, and electric autos.

Miller, W. N.↗

Execution time supports for adaptive scientific algorithms on distributed memory machines

Optimizations are considered that are required for efficient execution of code segments that consists of loops over distributed data structures. The PARTI (Parallel Automated Runtime Toolkit at ICASE) execution time primitives are designed to carry out these optimizations and can be used to implement a wide range of scientific algorithms on distributed memory machines. These primitives allow the user to control array mappings in a way that gives an appearance of shared memory. Computations can be based on a global index set. Primitives are used to carry out gather and scatter operations on distributed arrays. Communications patterns are derived at runtime, and the appropriate send and receive messages are automatically generated.

Berryman, Harry↗

Geoid and topography for infinite Prandtl number convection in a spherical shell

Geoid anomalies and surface and lower-boundary topographies are calculated for numerically generated thermal convection for an infinite Prandtl number, Boussinesq, axisymmetric spherical fluid shell with constant gravity and viscosity, for heating both entirely from below and entirely from within. Convection solutions are obtained for Rayleigh numbers Ra up to 20 times the critical Ra in heating from below and 27 times critical for heating from within. Geoid parallels surface undulations, and boundary deformation generally increases with increasing cell wavelength. Dimensionless geoid and topography in heating from below are about 5 times greater than in heating from within. Values for heating from within correlate more closely with geophysical data than values from heating from below, suggesting a predominance of internal heating in the mantle. The study emphasizes that dynamically induced topography and geoid are sensitive to the mode of heating in the earth's mantle.

Bercovici, D.↗

Real-Time Bayesian Inference at Extreme Scale: A Digital Twin for Tsunami Early Warning Applied to the Cascadia Subduction Zone

We present a Bayesian inversion-based digital twin that employs acoustic pressure data from seafloor sensors, along with 3D coupled acoustic–gravity wave equations, to infer earthquake-induced spatiotemporal seafloor motion in real time and forecast tsunami propagation toward coastlines for early warning with quantified uncertainties. Our target is the Cascadia subduction zone, with one billion parameters. Computing the posterior mean alone would require 50 years on a 512 GPU machine. Instead, exploiting the shift invariance of the parameter-to-observable map and devising novel parallel algorithms, we induce a fast offline–online decomposition. The offline component requires just one adjoint wave propagation per sensor; using MFEM, we scale this part of the computation to the full El Capitan system (43,520 GPUs) with 92% weak parallel efficiency. Moreover, given real-time data, the online component exactly solves the Bayesian inverse and forecasting problems in 0.2 seconds on a modest GPU system, a ten-billion-fold speedup.

97 MATHEMATICS AND COMPUTING↗

Parallel-Processing Software for Creating Mosaic Images

A computer program implements parallel processing for nearly real-time creation of panoramic mosaics of images of terrain acquired by video cameras on an exploratory robotic vehicle (e.g., a Mars rover). Because the original images are typically acquired at various camera positions and orientations, it is necessary to warp the images into the reference frame of the mosaic before stitching them together to create the mosaic. [Also see "Parallel-Processing Software for Correlating Stereo Images," Software Supplement to NASA Tech Briefs, Vol. 31, No. 9 (September 2007) page 26.] The warping algorithm in this computer program reflects the considerations that (1) for every pixel in the desired final mosaic, a good corresponding point must be found in one or more of the original images and (2) for this purpose, one needs a good mathematical model of the cameras and a good correlation of individual pixels with respect to their positions in three dimensions. The desired mosaic is divided into slices, each of which is assigned to one of a number of central processing units (CPUs) operating simultaneously. The results from the CPUs are gathered and placed into the final mosaic. The time taken to create the mosaic depends upon the number of CPUs, the speed of each CPU, and whether a local or a remote data-staging mechanism is used.

Klimeck, Gerhard↗

Development of Fast Algorithms Using Recursion, Nesting and Iterations for Computational Electromagnetics

In the first phase of our work, we have concentrated on laying the foundation to develop fast algorithms, including the use of recursive structure like the recursive aggregate interaction matrix algorithm (RAIMA), the nested equivalence principle algorithm (NEPAL), the ray-propagation fast multipole algorithm (RPFMA), and the multi-level fast multipole algorithm (MLFMA). We have also investigated the use of curvilinear patches to build a basic method of moments code where these acceleration techniques can be used later. In the second phase, which is mainly reported on here, we have concentrated on implementing three-dimensional NEPAL on a massively parallel machine, the Connection Machine CM-5, and have been able to obtain some 3D scattering results. In order to understand the parallelization of codes on the Connection Machine, we have also studied the parallelization of 3D finite-difference time-domain (FDTD) code with PML material absorbing boundary condition (ABC). We found that simple algorithms like the FDTD with material ABC can be parallelized very well allowing us to solve within a minute a problem of over a million nodes. In addition, we have studied the use of the fast multipole method and the ray-propagation fast multipole algorithm to expedite matrix-vector multiplication in a conjugate-gradient solution to integral equations of scattering. We find that these methods are faster than LU decomposition for one incident angle, but are slower than LU decomposition when many incident angles are needed as in the monostatic RCS calculations.

Chew, W. C.↗

The PISCES 2 parallel programming environment

PISCES 2 is a programming environment for scientific and engineering computations on MIMD parallel computers. It is currently implemented on a flexible FLEX/32 at NASA Langley, a 20 processor machine with both shared and local memories. The environment provides an extended Fortran for applications programming, a configuration environment for setting up a run on the parallel machine, and a run-time environment for monitoring and controlling program execution. This paper describes the overall design of the system and its implementation on the FLEX/32. Emphasis is placed on several novel aspects of the design: the use of a carefully defined virtual machine, programmer control of the mapping of virtual machine to actual hardware, forces for medium-granularity parallelism, and windows for parallel distribution of data. Some preliminary measurements of storage use are included.

Pratt, Terrence W.↗

Parallel Wavefront Analysis for a 4D Interferometer

This software provides a programming interface for automating data collection with a PhaseCam interferometer from 4D Technology, and distributing the image-processing algorithm across a cluster of general-purpose computers. Multiple instances of 4Sight (4D Technology s proprietary software) run on a networked cluster of computers. Each connects to a single server (the controller) and waits for instructions. The controller directs the interferometer to several images, then assigns each image to a different computer for processing. When the image processing is finished, the server directs one of the computers to collate and combine the processed images, saving the resulting measurement in a file on a disk. The available software captures approximately 100 images and analyzes them immediately. This software separates the capture and analysis processes, so that analysis can be done at a different time and faster by running the algorithm in parallel across several processors. The PhaseCam family of interferometers can measure an optical system in milliseconds, but it takes many seconds to process the data so that it is usable. In characterizing an adaptive optics system, like the next generation of astronomical observatories, thousands of measurements are required, and the processing time quickly becomes excessive. A programming interface distributes data processing for a PhaseCam interferometer across a Windows computing cluster. A scriptable controller program coordinates data acquisition from the interferometer, storage on networked hard disks, and parallel processing. Idle time of the interferometer is minimized. This architecture is implemented in Python and JavaScript, and may be altered to fit a customer s needs.

Rao, Shanti R.↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

An Accelerated Clip Algorithm for Unstructured Meshes: A Batch-Driven Approach

The clip technique is a popular method for visualizing complex structures and phenomena within 3D unstructured meshes. Meshes can be clipped by specifying a scalar isovalue to produce an output unstructured mesh with its external surface as the isovalue. Similar to isocontouring, the clipping process relies on scalar data associated with the mesh points, including scalar data generated by implicit functions such as planes, boxes, and spheres, which facilitates the visualization of results interior to the grid. In this paper, we introduce a novel batch-driven parallel algorithm based on a sequential clip algorithm designed for high-quality results in partial volume extraction. Our algorithm comprises five passes, each progressively processing data to generate the resulting clipped unstructured mesh. The novelty lies in the use of fixed-size batches of points and cells, which enable rapid workload trimming and parallel processing, leading to a significantly improved memory footprint and run-time performance compared to the original version. On a 32-core CPU, the proposed batch-driven parallel algorithm demonstrates a run-time speed-up of up to 32.6x and a memory footprint reduction of up to 4.37x compared to the existing sequential algorithm. The software is currently available under an open-source license in the VTK visualization system.

Tsalikis, Spiros↗

Thermal Field Imaging Using Ultrasound

It is often desirable to be able to determine the temperature field in the interiors of opaque fluids forced into convection by externally imposed temperature gradients. To measure the temperature at a point in an opaque fluid in the usual fashion requires insertion of a probe, and to determine the full field therefore requires either the ability to move this probe or the introduction of multiple probes. Neither of these solutions is particularly satisfactory, although they can lead to quite accurate measurements. As an alternative we have investigated the use of ultrasound as a relatively non-intrusive probe of the temperature field in convecting opaque fluids. The temperature dependence of the sound velocity can be sufficiently great to permit a determination of the temperature from timing the traversal of an ultrasound pulse across a chamber. In this paper we will present our results on convecting flows of transparent and opaque fluids. Our experimental cells consist of relatively narrow rectangular cavities made of thermally insulating materials on the sides, and metal top and bottom plates. The ultrasound transducer is powered by a pulser/receiver, the signal output of which goes to a very high speed signal averager. The average of several hundred to several thousand signals is then sent to a computer for storage and analysis. The experimental procedure is to establish a convective flow by imposing a vertical temperature gradient on the chamber, and then to measure, at several regularly spaced locations, the transit time for an ultrasound pulse to traverse the chamber horizontally (parallel to the convecting rolls) and return to the transducer. The transit time is related to the temperature of the fluid through which the sound pulse travels. Knowing the relationship between transit time and temperature (determined in a separate experiment), we can extract the average temperature across the chamber at that location. By changing the location of the transducer it is then possible to find the average temperature at different locations along the chamber, thereby determining the temperature profile along the system. (In the future we will construct an array of transducers. This will give us the capability to determine the temperature profile much more rapidly than at present, an important consideration if time-dependent phenomena are to be studied.) To validate our procedure we introduced encapsulated liquid crystal particles into glycerol. The liquid crystal particles' color varies depending on the temperature of the fluid. A photograph of the fluid through transparent sidewalls therefore gives a picture of the temperature field of the convecting fluid, independent of our ultrasound imaging. A representative result is shown in the Figure 1, which reveals a very satisfying correspondence between the two techniques. Therefore we have a great deal of confidence that the ultrasound imaging approach is indeed measuring the actual temperature profile of the fluid. The technique has also been applied to convecting liquid metal flows, and representative data will be presented from those experiments as well.

Andereck, D.↗