Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68

Scheduling for Parallel Supercomputing: A Historical Perspective of Achievable Utilization

The NAS facility has operated parallel supercomputers for the past 11 years, including the Intel iPSC/860, Intel Paragon, Thinking Machines CM-5, IBM SP-2, and Cray Origin 2000. Across this wide variety of machine architectures, across a span of 10 years, across a large number of different users, and through thousands of minor configuration and policy changes, the utilization of these machines shows three general trends: (1) scheduling using a naive FIFO first-fit policy results in 40-60% utilization, (2) switching to the more sophisticated dynamic backfilling scheduling algorithm improves utilization by about 15 percentage points (yielding about 70% utilization), and (3) reducing the maximum allowable job size further increases utilization. Most surprising is the consistency of these trends. Over the lifetime of the NAS parallel systems, we made hundreds, perhaps thousands, of small changes to hardware, software, and policy, yet, utilization was affected little. In particular these results show that the goal of achieving near 100% utilization while supporting a real parallel supercomputing workload is unrealistic.

Jones, James Patton↗

Quantum simulation of excited states from parallel contracted quantum eigensolvers

Abstract Computing excited-state properties of molecules and solids is considered one of the most important near-term applications of quantum computers. While many of the current excited-state quantum algorithms differ in circuit architecture, specific exploitation of quantum advantage, or result quality, one common feature is their rooting in the Schrödinger equation. However, through contracting (or projecting) the eigenvalue equation, more efficient strategies can be designed for near-term quantum devices. Here we demonstrate that when combined with the Rayleigh–Ritz variational principle for mixed quantum states, the ground-state contracted quantum eigensolver (CQE) can be generalized to compute any number of quantum eigenstates simultaneously. We introduce two excited-state (anti-Hermitian) CQEs that perform the excited-state calculation while inheriting many of the remarkable features of the original ground-state version of the algorithm, such as its scalability. To showcase our approach, we study several model and chemical Hamiltonians and investigate the performance of different implementations.

Physics↗

Field lines and magnetic surfaces in a two-component slab/2D model of interplanetary magnetic fluctuations

A two-component model for the spectrum of interplanetary magnetic fluctuations was proposed on the basis of ISEE observations, and has found an intriguing level of application in other solar wind studies. The model fluctuations consist of a fraction of 'slab' fluctuations, varying only in the direction parallel to the locally uniform mean magnetic field B(0) and a complement of 2D (two-dimensional) fluctuations that vary in the directions transverse to B(0). We have developed an spectral method computational algorithm for computing the magnetic flux surfaces (flux tubes) associated with the composite model, based upon a precise analogy with equations for ideal transport of a passive scalar in planar two dimensional geometry. Visualization of various composite models will be presented, including the 80 percent 2D/ 20 percent slab model with delta B/B(0) approximately equals 1 and a minus 5/3 spectral law, that is thought to approximately represent a snapshot of solar wind turbulence. Characteristically, the visualizations show that flux tubes, even when defined as regular on some plane, shred and disperse rapidly as they are viewed along the parallel direction. This diffusive process, which generalizes the standard picture of field line random walk, will be discussed in detail. Evidently, the traditional picture that flux tubes randomize like strands of spaghetti with a uniform tangle along the axial direction is in need of modification.

Matthaeus, W. H.↗

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Contextual subspace variational quantum eigensolver calculation of the dissociation curve of molecular nitrogen on a superconducting quantum computer

Abstract We present an experimental demonstration of the Contextual Subspace Variational Quantum Eigensolver on superconducting hardware. Calculating the potential energy curve of molecular nitrogen proves challenging for many conventional quantum chemistry techniques, since static correlation dominates in the dissociation limit. Our quantum simulations retain good agreement with the Full Configuration Interaction energy, outperforming all benchmarked single-reference wavefunction techniques in capturing the bond-breaking appropriately. Moreover, our methodology is competitive with multiconfigurational approaches but at a saving of quantum resource, meaning larger active spaces can be treated for a fixed qubit allowance. To achieve this result, we deploy an error mitigation/suppression strategy comprised of Dynamical Decoupling, Measurement-Error Mitigation and Zero-Noise Extrapolation. Circuit parallelization also provides passive noise-averaging and improves the effective shot yield to reduce the measurement overhead. Furthermore, we introduce a modified adaptive ansatz construction algorithm that incorporates hardware awareness into our variational circuits, minimizing the transpilation cost for the target qubit topology.

Physics↗

Asynchronous multilevel adaptive methods for solving partial differential equations on multiprocessors - Performance results

The fast adaptive composite grid method (FAC) is an algorithm that uses various levels of uniform grids (global and local) to provide adaptive resolution and fast solution of PDEs. Like all such methods, it offers parallelism by using possibly many disconnected patches per level, but is hindered by the need to handle these levels sequentially. The finest levels must therefore wait for processing to be essentially completed on all the coarser ones. A recently developed asynchronous version of FAC, called AFAC, completely eliminates this bottleneck to parallelism. This paper describes timing results for AFAC, coupled with a simple load balancing scheme, applied to the solution of elliptic PDEs on an Intel iPSC hypercube. These tests include performance of certain processes necessary in adaptive methods, including moving grids and changing refinement. A companion paper reports on numerical and analytical results for estimating convergence factors of AFAC applied to very large scale examples.

Mccormick, S.↗

Unsteady turbomachinery flow simulations on massively parallel architectures

The accurate numerical simulation of unsteady, three-dimensional viscous flow in turbomachines is computationally very intensive, requiring prohibitively large amounts of computer time on current vector supercomputers. In recent years, computer systems based on massively parallel architectures have been developed that offer the promise of meeting the computational power requirements of such large-scale simulations. However, a rethinking of existing algorithms and methodology is required in order to fully harness the computational power of such architectures. In this paper the capabilities of the Connection Machine (CM-2) in predicting unsteady flows in turbomachines are evaluated. The implementation on the CM-2 of an implicit, time-accurate, zonal algorithm for the Navier-Stokes equations in two dimensions is described. Programming issues and modifications made to the original algorithm (developed for vector, pipelined supercomputers) in order to improve performance on the CM-2 are outlined. Algorithm performance is evaluated and compared with a functionally equivalent code for the CRAY-YMP.

Madavan, N. K.↗

On Improving Efficiency of Differential Evolution for Aerodynamic Shape Optimization Applications

Differential Evolution (DE) is a simple and robust evolutionary strategy that has been provEn effective in determining the global optimum for several difficult optimization problems. Although DE offers several advantages over traditional optimization approaches, its use in applications such as aerodynamic shape optimization where the objective function evaluations are computationally expensive is limited by the large number of function evaluations often required. In this paper various approaches for improving the efficiency of DE are reviewed and discussed. Several approaches that have proven effective for other evolutionary algorithms are modified and implemented in a DE-based aerodynamic shape optimization method that uses a Navier-Stokes solver for the objective function evaluations. Parallelization techniques on distributed computers are used to reduce turnaround times. Results are presented for standard test optimization problems and for the inverse design of a turbine airfoil. The efficiency improvements achieved by the different approaches are evaluated and compared.

Madavan, Nateri K.↗

High-speed computerized tomography

The development of a high-speed reconstruction processor and a channelized architecture to use with a high-resolution tomographic unit is discussed with attention to the convolution reconstruction algorithm. By means of this algorithm, input data and intermediate result precision required throughout the algorithm execution have been studied with computer simulation using profile data derived from mathematically simulated test objects and experimental animal data. A prototype section for a highly parallel all-digital system executes 60 million arithmetic operations per second, and the full-scale version is expected to reconstruct 500 to 1000 cross sections per second.

Swartzlander, E. E., Jr.↗

Cumulative reports and publications through 31 December 1983

All reports for the calendar years 1975 through December 1983 are listed by author. Since ICASE reports are intended to be preprints of articles for journals and conference proceedings, the published reference is included when available. Thirteen older journal and conference proceedings references are included as well as five additional reports by ICASE personnel. Major categories of research covered include: (1) numerical methods, with particular emphasis on the development and analysis of basic algorithms; (2) computational problems in engineering and the physical sciences, particularly fluid dynamics, acoustics, structural analysis, and chemistry; and (3) computer systems and software, especially vector and parallel computers, microcomputers, and data management.

Source record↗

Efficiency of group implicit concurrent algorithms for transient finite element analysis

The performance of group implicit algorithms is assessed on actual concurrent computers. It is shown that, as the number of subdomains is increased, performance enhancements are derived from two sources: the increased parallelism in the computations; and a reduction in equation solving effort. Moreover, these two performance enhancements are synergistic, in the sense that the corresponding speed-ups are multiplied, rather than merely added. Simulations on a 32-node hypercube are presented for which the interprocessor communications efficiencies obtained are consistently in excess of 90 percent.

Ortiz, M.↗

Design of optimal correlation filters for hybrid vision systems

Research is underway at the NASA Johnson Space Center on the development of vision systems that recognize objects and estimate their position by processing their images. This is a crucial task in many space applications such as autonomous landing on Mars sites, satellite inspection and repair, and docking of space shuttle and space station. Currently available algorithms and hardware are too slow to be suitable for these tasks. Electronic digital hardware exhibits superior performance in computing and control; however, they take too much time to carry out important signal processing operations such as Fourier transformation of image data and calculation of correlation between two images. Fortunately, because of the inherent parallelism, optical devices can carry out these operations very fast, although they are not quite suitable for computation and control type operations. Hence, investigations are currently being conducted on the development of hybrid vision systems that utilize both optical techniques and digital processing jointly to carry out the object recognition tasks in real time. Algorithms for the design of optimal filters for use in hybrid vision systems were developed. Specifically, an algorithm was developed for the design of real-valued frequency plane correlation filters. Furthermore, research was also conducted on designing correlation filters optimal in the sense of providing maximum signal-to-nose ratio when noise is present in the detectors in the correlation plane. Algorithms were developed for the design of different types of optimal filters: complex filters, real-value filters, phase-only filters, ternary-valued filters, coupled filters. This report presents some of these algorithms in detail along with their derivations.

Rajan, Periasamy K.↗

Execution time support for scientific programs on distributed memory machines

Optimizations are considered that are required for efficient execution of code segments that consists of loops over distributed data structures. The PARTI (Parallel Automated Runtime Toolkit at ICASE) execution time primitives are designed to carry out these optimizations and can be used to implement a wide range of scientific algorithms on distributed memory machines. These primitives allow the user to control array mappings in a way that gives an appearance of shared memory. Computations can be based on a global index set. Primitives are used to carry out gather and scatter operations on distributed arrays. Communications patterns are derived at runtime, and the appropriate send and receive messages are automatically generated.

Berryman, Harry↗

Ionospheric refraction effects on TOPEX orbit determination accuracy using the Tracking and Data Relay Satellite System (TDRSS)

This investigation concerns the effects on Ocean Topography Experiment (TOPEX) spacecraft operational orbit determination of ionospheric refraction error affecting tracking measurements from the Tracking and Data Relay Satellite System (TDRSS). Although tracking error from this source is mitigated by the high frequencies (K-band) used for the space-to-ground links and by the high altitudes for the space-to-space links, these effects are of concern for the relatively high-altitude (1334 kilometers) TOPEX mission. This concern is due to the accuracy required for operational orbit-determination by the Goddard Space Flight Center (GSFC) and to the expectation that solar activity will still be relatively high at TOPEX launch in mid-1992. The ionospheric refraction error on S-band space-to-space links was calculated by a prototype observation-correction algorithm using the Bent model of ionosphere electron densities implemented in the context of the Goddard Trajectory Determination System (GTDS). Orbit determination error was evaluated by comparing parallel TOPEX orbit solutions, applying and omitting the correction, using the same simulated TDRSS tracking observations. The tracking scenarios simulated those planned for the observation phase of the TOPEX mission, with a preponderance of one-way return-link Doppler measurements. The results of the analysis showed most TOPEX operational accuracy requirements to be little affected by space-to-space ionospheric error. The determination of along-track velocity changes after ground-track adjustment maneuvers, however, is significantly affected when compared with the stringent 0.1-millimeter-per-second accuracy requirements, assuming uncoupled premaneuver and postmaneuver orbit determination. Space-to-space ionospheric refraction on the 24-hour postmaneuver arc alone causes 0.2 millimeter-per-second errors in along-track delta-v determination using uncoupled solutions. Coupling the premaneuver and postmaneuver solutions, however, appears likely to reduce this figure substantially. Plans and recommendations for response to these findings are presented.

Radomski, M. S.↗

An Efficient GPU-Accelerated Multi-Source Global Fit Pipeline for LISA Data Analysis

The large-scale analysis task of deciphering gravitational wave signals in the LISA data stream will be difficult, requiring a large amount of computational resources and extensive development of computational methods. Its high dimensionality, multiple model types, and complicated noise profile require a global fit to all parameters and input models simultaneously. In this work, we detail our global fit algorithm, called “Erebor,” designed to accomplish this challenging task. It is capable of analysing current state-of-the-art datasets and then growing into the future as more pieces of the pipeline are completed and added. We describe our pipeline strategy, the algorithmic setup, and the results from our analysis of the LDC2A Sangria dataset, which contains Massive Black Hole Binaries, compact Galactic Binaries, and a parameterized noise spectrum whose parameters are unknown to the user. The Erebor algorithm includes three unique and very useful contributions: GPU acceleration for enhanced computational efficiency; ensemble MCMC sampling with multiple MCMC walkers per temperature for better mixing and parallelized sample creation; and special online updates to reversible-jump (or trans-dimensional) sampling distributions to ensure sampler mixing and accurate initial estimates for detectable sources in the data. We recover posterior distributions for all 15 (6) of the injected MBHBs in the LDC2A training (hidden) dataset. We catalog ∼12000 Galactic Binaries (∼8000 as high confidence detections) for both the training and hidden datasets. All of the sources and their posterior distributions are provided in publicly available catalogs.

LISA global fit↗

Exploiting Modern C++ for Portable Parallel Programming in Lattice QCD Applications

The evolution of ISO C++ standards increasingly serves the needs of scientific computing, offering potential benefits for developing portable applications. The recent revisions of C++ programming language, for instance, introduces a suite of algorithms capable of being executed on accelerators. Although this approach may not yield best performance, it can present a viable balance between code productivity and computational efficiency. In this report, we discuss the implementation of the HISQ operator utilizing a range of features from the C++17/20/23 standards and include an assessment of their performance.

Strelchenko, Alexei↗

Some Experiences with Nonoverlapping Schur Complement Parallel Preconditioning for CFD Calculations

In this work we consider solving matrices which arise from the discretization of advection-diffusion field equations on arbitrary triangulated domains using stabilized numerical methods. The talk will discuss several candidate matrix preconditioning algorithms based on the 2 x 2 block factorization induced by an apriori partitioning of the triangulated domain. Application of the 2 x 2 block preconditioner requires the formation and inversion of the Schur complement submatrix. We consider several strategies for simplifying this task: incomplete Schur complement factorizations, drop tolerance element filling, Schur complement probing, and localized Schur complement inversion. Numerical results will be shown comparing performance and efficiency of these approximations. The matrix preconditioner has also been embedded into a Newton algorithm for solving the nonlinear Euler and Navier-Stokes equations governing compressible flow. The remainder of the talk will show numerous examples in CFD to demonstrate the efficiency and robustness of the techniques.

Barth, Timothy J.↗

Fast Particle Methods for Multiscale Phenomena Simulations

We are developing particle methods oriented at improving computational modeling capabilities of multiscale physical phenomena in : (i) high Reynolds number unsteady vortical flows, (ii) particle laden and interfacial flows, (iii)molecular dynamics studies of nanoscale droplets and studies of the structure, functions, and evolution of the earliest living cell. The unifying computational approach involves particle methods implemented in parallel computer architectures. The inherent adaptivity, robustness and efficiency of particle methods makes them a multidisciplinary computational tool capable of bridging the gap of micro-scale and continuum flow simulations. Using efficient tree data structures, multipole expansion algorithms, and improved particle-grid interpolation, particle methods allow for simulations using millions of computational elements, making possible the resolution of a wide range of length and time scales of these important physical phenomena.The current challenges in these simulations are in : [i] the proper formulation of particle methods in the molecular and continuous level for the discretization of the governing equations [ii] the resolution of the wide range of time and length scales governing the phenomena under investigation. [iii] the minimization of numerical artifacts that may interfere with the physics of the systems under consideration. [iv] the parallelization of processes such as tree traversal and grid-particle interpolations We are conducting simulations using vortex methods, molecular dynamics and smooth particle hydrodynamics, exploiting their unifying concepts such as : the solution of the N-body problem in parallel computers, highly accurate particle-particle and grid-particle interpolations, parallel FFT's and the formulation of processes such as diffusion in the context of particle methods. This approach enables us to transcend among seemingly unrelated areas of research.

Koumoutsakos, P.↗