Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70

Improved Finite-Volume Method for Radiative Hydrodynamics

Fully coupled simulations of hydrodynamics and radiative transfer are essential to a number of fields ranging from astrophysics to engineering applications. Of particular interest in this work are hypersonic atmospheric entries and associated experimental apparatus, e.g., shock tubes and high enthalpy testing facilities. The radiative transfer calculations must supply to the CFD a heating term in the energy equation in the form of the divergence of the radiative heat flux and the radiative heat fluxes to bounding surfaces. It is most efficient to solve the radiative transfer equation on the same grid as the CFD solution, and this work presents an algorithm with improved accuracy for such simulations on structured and unstructured grids compared to more conventional approaches. Results will be shown for shock radiation during hypersonic reentry. Issues of parallelization within a radiation sweep will also be discussed.

Wray, Alan↗

The Use of Field Programmable Gate Arrays (FPGA) in Small Satellite Communication Systems

This paper will describe the use of digital Field Programmable Gate Arrays (FPGA) to contribute to advancing the state-of-the-art in software defined radio (SDR) transponder design for the emerging SmallSat and CubeSat industry and to provide advances for NASA as described in the TAO5 Communication and Navigation Roadmap (Ref 4). The use of software defined radios (SDR) has been around for a long time. A typical implementation of the SDR is to use a processor and write software to implement all the functions of filtering, carrier recovery, error correction, framing etc. Even with modern high speed and low power digital signal processors, high speed memories, and efficient coding, the compute intensive nature of digital filters, error correcting and other algorithms is too much for modern processors to get efficient use of the available bandwidth to the ground. By using FPGAs, these compute intensive tasks can be done in parallel, pipelined fashion and more efficiently use every clock cycle to significantly increase throughput while maintaining low power. These methods will implement digital radios with significant data rates in the X and Ka bands. Using these state-of-the-art technologies, unprecedented uplink and downlink capabilities can be achieved in a 1/2 U sized telemetry system. Additionally, modern FPGAs have embedded processing systems, such as ARM cores, integrated inside the FPGA allowing mundane tasks such as parameter commanding to occur easily and flexibly. Potential partners include other NASA centers, industry and the DOD. These assets are associated with small satellite demonstration flights, LEO and deep space applications. MSFC currently has an SDR transponder test-bed using Hardware-in-the-Loop techniques to evaluate and improve SDR technologies.

Varnavas, Kosta↗

Development of the Science Data System for the International Space Station Cold Atom Lab

Cold Atom Laboratory (CAL) is a facility that will enable scientists to study ultra-cold quantum gases in a microgravity environment on the International Space Station (ISS) beginning in 2016. The primary science data for each experiment consists of two images taken in quick succession. The first image is of the trapped cold atoms and the second image is of the background. The two images are subtracted to obtain optical density. These raw Level 0 atom and background images are processed into the Level 1 optical density data product, and then into the Level 2 data products: atom number, Magneto-Optical Trap (MOT) lifetime, magnetic chip-trap atom lifetime, and condensate fraction. These products can also be used as diagnostics of the instrument health. With experiments being conducted for 8 hours every day, the amount of data being generated poses many technical challenges, such as downlinking and managing the required data volume. A parallel processing design is described, implemented, and benchmarked. In addition to optimizing the data pipeline, accuracy and speed in producing the Level 1 and 2 data products is key. Algorithms for feature recognition are explored, facilitating image cropping and accurate atom number calculations.

bose einstein condensate↗

Simulations of particle acceleration in parallel shocks: Direct comparison between Monte Carlo and one-dimensional hybrid codes

We have made a direct comparison between two different computer simulations of a plane, parallel, collisionless shock including particle acceleration to energies typical of those of diffuse ions observed at the earth bow shock. Despite the fact that the one-dimensional hybrid and Monte Carlo techniques employ entirely different algorithms, they give surprisingly close agreement in the overall shapes of the complete distribution functions for protons as well as heavier ions. Both methods show that energetic ions emerge smoothly from the background thermal plasma with approximately the same relative injection rate and that the fraction of the incoming plasma's energy flux that is converted into downstream enthalpy flux of the accelerated population (i.e., the acceleration efficiency) is similar in the two cases. The fraction of the downstream proton distribution made up of superthermal particles is quite large, with at least 10% of the energy flux going into protons with energies above 10 keV. In addition, an upstream precursor, produced by backstreaming energetic particles, is present in both shocks, although the Monte Carlo precursor is considerably longer than that produced in the hybrid shock. These results offer convincing evidence that, at least in these ways, the two simulations are consistent in their description of parallel shock structure and particle acceleration, and they lay the groundwork for development of shock models employing a combination of both methods.

Ellison, Donald C.↗

Improved Characterization of PSC Processes Derived from a Third-Generation CALIOP and MLS Detection and Composition Classification Algorithm

The new 3-year CloudSat and CALIPSO Science Team project described in this poster will use a unique combination of data from the Cloud-Aerosol LIdar with Orthogonal Polarization (CALIOP) instrument on CALIPSO and the Microwave Limb Sounder (MLS) on Aura, in conjunction with supporting meteorological information and detailed modeling studies, to advance our understanding of polar stratospheric cloud (PSC) processes and their role in ozone depletion. We will develop a third-generation (Gen3) PSC detection and composition algorithm that incorporates a new, more robust two-dimensional, multi-channel CALIOP feature detection scheme (2D-McDA). We will also devise and implement an improved two-dimensional PSC composition classification scheme that utilizes multiple parameters (e.g., CALIOP 532-nm parallel and perpendicular scattering ratios, CALIOP 1064-nm total scattering ratio, MLS HNO3 and H2O, ambient temperature, and temperature histories) in a Bayesian approach to determine the most likely PSC composition and help constrain solid PSC particle number density and size/shape. The combined CALIOP/MLS analyses will allow us to study in detail the full life cycle of PSCs and their resulting impact on gas-phase HNO3 and H2O, which should lead to improved parameterizations of PSC microphysics in global CCMs where detailed particle information is not available. The Gen3 CALIOP PSC algorithm will be a natural stepping-stone toward the analysis of data collected during future spaceborne lidar missions, such as NASA’s Atmosphere Observation System (AtmOS) mission currently scheduled for launch late in this decade. We will also investigate possible trends in PSC occurrence and composition over the entire CALIOP data record and through further comparisons with the Stratospheric Aerosol Measurement (SAM) II solar occultation PSC record from 1979-1989. Finally, we will validate the mountain-wave parameterization and PSC schemes used in the UM-UKCA (Unified Model coupled to the United Kingdom Chemistry and Aerosol module) chemistry-climate model through detailed comparisons with earlier CALIOP PSC data products and those developed under this proposal.

CALIPSO↗

Improved Characterization of PSC Processes Derived from a Third-Generation CALIOP and MLS Detection and Composition Classification Algorithm

The new 3-year CloudSat and CALIPSO Science Team project described in this poster will use a unique combination of data from the Cloud-Aerosol LIdar with Orthogonal Polarization (CALIOP) instrument on CALIPSO and the Microwave Limb Sounder (MLS) on Aura, in conjunction with supporting meteorological information and detailed modeling studies, to advance our understanding of polar stratospheric cloud (PSC) processes and their role in ozone depletion. We will develop a third-generation (Gen3) PSC detection and composition algorithm that incorporates a new, more robust two-dimensional, multi-channel CALIOP feature detection scheme (2D-McDA). We will also devise and implement an improved two-dimensional PSC composition classification scheme that utilizes multiple parameters (e.g., CALIOP 532-nm parallel and perpendicular scattering ratios, CALIOP 1064-nm total scattering ratio, MLS HNO3 and H2O, ambient temperature, and temperature histories) in a Bayesian approach to determine the most likely PSC composition and help constrain solid PSC particle number density and size/shape. The combined CALIOP/MLS analyses will allow us to study in detail the full life cycle of PSCs and their resulting impact on gas-phase HNO3 and H2O, which should lead to improved parameterizations of PSC microphysics in global chemistry-climate models (CCMs) where detailed particle information is not available. The Gen3 CALIOP PSC algorithm will be a natural stepping-stone toward the analysis of data collected during future spaceborne lidar missions, such as NASA’s Atmosphere Observation System (AOS) mission currently scheduled for launch late in this decade. We will also investigate possible trends in PSC occurrence and composition over the entire CALIOP data record and through further comparisons with the Stratospheric Aerosol Measurement (SAM) II solar occultation PSC record from 1979-1989. Finally, we will validate the mountain-wave parameterization and PSC schemes used in the UM-UKCA (Unified Model coupled to the United Kingdom Chemistry and Aerosol module) CCM through detailed comparisons with earlier CALIOP PSC data products and those developed under this proposal.

CALIPSO↗

Monte Carlo Explicitly Correlated Second-Order Many-Body Green’s Function Calculations of Semiconductor Band Gaps

A systematically converging series of ab initio, post-density-functional, size-consistent, electron-correlated approximations is desired for predictive computing of felectronic band structures of insulating, semiconducting, and metallic solids. A series that meets all of these desiderata (except the applicability to metals) is ab initio many-body Green's function theory based on Gaussian-type-orbital (GTO) basis sets. Here, its leading-order approximation, the second-order Green's function (GF2) method in the diagonal and frequency-independent approximations with the aug-cc-pVDZ basis set, is applied to the fundamental band gaps of three semiconductors (diamond, silicon, and silicon carbide in the zincblende structure) using cluster models. Corrections are made to the basis-set-incompleteness errors by the explicit-correlation (F12) ansatz (GF2-F12) for the valence band edges. The crystals are modeled as surface-passivated clusters of increasing sizes, whose wave functions are expanded by up to 2709 GTO basis functions. Immense computational costs of these calculations are overcome by the highly scalable stochastic algorithm of the Monte Carlo GF2-F12 method, whose operation cost per state increases only as a cubic power of system size, which has a tiny memory footprint and easily achieves near-perfect parallel efficiency on thousands of CPUs or on hundreds of GPUs. The correlated, F12-corrected highest-occupied and lowest-unoccupied molecular-orbital energy (HOMO-LUMO) gap is 5.78 ± 0.07 eV for C 87 H 76 as compared with the experimental value of the fundamental (indirect) band gap of bulk diamond at 5.48 eV. The correlated, F12-corrected HOMO-LUMO gaps for Si 75 H 76 and Si 32 C 43 H 76 are 2.56 ± 0.15 eV and 3.50 ± 0.12 eV, respectively, which are expected to decrease further with increasing cluster sizes. As a result, the experimental fundamental (indirect) band gaps of bulk silicon and silicon carbide are 1.17 eV and 2.42 eV, respectively.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Portable HCAL reconstruction in the CMS detector using the Alpaka library

CMS has deployed a number of different GPU algorithms at the High-Level Trigger (HLT) in Run 3. As the code base for GPU algorithms continues to grow, the burden for developing and maintaining separate implementations for GPU and CPU becomes increasingly challenging. To mitigate this, CMS has adopted the Alpaka (Abstraction Library for Parallel Kernel Acceleration) library as the performance portability solution to provide a single-code base for parallel execution on both GPUs and CPUs in CMS software (CMSSW). A direct CUDA version of HCAL energy reconstruction, called Minimization At Hcal, Iteratively (MAHI), has been deployed at the HLT in the 2022-2023 data taking period. This contribution will describe how the CUDA version is converted into a portable implementation using the Alpaka library. We will discuss the porting experience from CUDA to Alpaka, the validation process and the performance of the Alpaka version in CPU and GPU.

Kwok, Martin↗

Solution of partial differential equations on vector and parallel computers

The present status of numerical methods for partial differential equations on vector and parallel computers was reviewed. The relevant aspects of these computers are discussed and a brief review of their development is included, with particular attention paid to those characteristics that influence algorithm selection. Both direct and iterative methods are given for elliptic equations as well as explicit and implicit methods for initial boundary value problems. The intent is to point out attractive methods as well as areas where this class of computer architecture cannot be fully utilized because of either hardware restrictions or the lack of adequate algorithms. Application areas utilizing these computers are briefly discussed.

Ortega, J. M.↗

Solution of partial differential equations on vector and parallel computers

The present status of numerical methods for partial differential equations on vector and parallel computers was reviewed. The relevant aspects of these computers are discussed and a brief review of their development is included, with particular attention paid to those characteristics that influence algorithm selection. Both direct and iterative methods are given for elliptic equations as well as explicit and implicit methods for initial boundary value problems. The intent is to point out attractive methods as well as areas where this class of computer architecture cannot be fully utilized because of either hardware restrictions or the lack of adequate algorithms. Application areas utilizing these computers are briefly discussed.

Ortega, J. M.↗

Hyperspectral data analysis procedures with reduced sensitivity to noise

Multispectral sensor systems have become steadily improved over the years in their ability to deliver increased spectral detail. With the advent of hyperspectral sensors, including imaging spectrometers, this technology is in the process of taking a large leap forward, thus providing the possibility of enabling delivery of much more detailed information. However, this direction of development has drawn even more attention to the matter of noise and other deleterious effects in the data, because reducing the fundamental limitations of spectral detail on information collection raises the limitations presented by noise to even greater importance. Much current effort in remote sensing research is thus being devoted to adjusting the data to mitigate the effects of noise and other deleterious effects. A parallel approach to the problem is to look for analysis approaches and procedures which have reduced sensitivity to such effects. We discuss some of the fundamental principles which define analysis algorithm characteristics providing such reduced sensitivity. One such analysis procedure including an example analysis of a data set is described, illustrating this effect.

Landgrebe, David A.↗

SAMPEX special pointing mode

A new pointing mode has been developed for the Solar, Anomalous, and Magnetospheric Particle Explorer (SAMPEX) spacecraft. This pointing mode orients the instrument boresights perpendicular to the field lines of the Earth's magnetic field in regions of low field strength and parallel to the field lines in regions of high field strength, to allow better characterization of heavy ions trapped by the field. The new mode uses magnetometer signals and is algorithmically simpler than the previous control mode, but it requires increased momentum wheel activity. It was conceived, designed, tested, coded, uplinked to the spacecraft, and activated in less than seven months.

Markley, F. Landis↗

Large-Scale Distributed Computational Fluid Dynamics on the Information Power Grid Using Globus

This paper describes an experiment in which a large-scale scientific application development for tightly-coupled parallel machines is adapted to the distributed execution environment of the Information Power Grid (IPG). A brief overview of the IPG and a description of the computational fluid dynamics (CFD) algorithm are given. The Globus metacomputing toolkit is used as the enabling device for the geographically-distributed computation. Modifications related to latency hiding and Load balancing were required for an efficient implementation of the CFD application in the IPG environment. Performance results on a pair of SGI Origin 2000 machines indicate that real scientific applications can be effectively implemented on the IPG; however, a significant amount of continued effort is required to make such an environment useful and accessible to scientists and engineers.

Barnard, Stephen↗

Performance Enhancements for the Lattice-Boltzmann Solver in the LAVA Framework

Performance enhancements in NASA's recently developed Lattice Boltzmann solver within the Launch Ascent and Vehicle Aerodynamics (LAVA) framework are presented. Two key algorithmic developments are highlighted. A coarse-fine interface treatment that discretely conserves mass and momentum has been implemented and successfully verified and validated. Code optimizations targeting improved serial and parallel performance were presented. For a simple turbulent Taylor-Green Vortex problem, we were able to demonstrate a 2.3 times speedup over the baseline code for a single Skylake-SP node containing 40 physical cores, and a 2.14 times speedup for 64 nodes containing 2560 physical cores. In addition, we were able to show that the optimizations enabled us to scale the code almost perfectly to 20480 physical cores where, including ghost cells, the problem size was 10 billion cells.

Barad, Michael↗

Performance Enhancements for the Lattice-Boltzmann Solver in the LAVA Framework

Performance enhancements in NASA's recently developed Lattice Boltzmann solver within the Launch Ascent and Vehicle Aerodynamics (LAVA) framework are presented. Two key algorithmic developments are highlighted. A coarse-fine interface treatment that discretely conserves mass and momentum has been implemented and successfully verified and validated. Code optimizations targeting improved serial and parallel performance were presented. For a simple turbulent Taylor-Green Vortex problem, we were able to demonstrate a 2.3 times speedup over the baseline code for a single Skylake-SP node containing 40 physical cores, and a 2.14 times speedup for 64 nodes containing 2560 physical cores. In addition, we were able to show that the optimizations enabled us to scale the code almost perfectly to 20480 physical cores where, including ghost cells, the problem size was 10 billion cells.

LAVA↗

Parallel-in-Time Solution of Scalar Nonlinear Conservation Laws

Here, we consider the parallel-in-time solution of scalar nonlinear conservation laws in one spatial dimension. The equations are discretized in space with a conservative finite-volume method using weighted essentially nonoscillatory (WENO) reconstructions, and in time with high-order explicit Runge–Kutta methods. The solution of the global, discretized space-time problem is sought via a nonlinear iteration that uses a novel linearization strategy in cases of nondifferentiable equations. Under certain choices of discretization and algorithmic parameters, the nonlinear iteration coincides with Newton’s method, although, more generally, it is a preconditioned residual correction scheme. At each nonlinear iteration, the linearized problem takes the form of a certain discretization of a linear conservation law over the space-time domain in question. An approximate parallel-in-time solution of the linearized problem is computed with a single multigrid reduction-in-time (MGRIT) iteration; however, any other effective parallel-in-time method could be used in its place. The MGRIT iteration employs a novel coarse-grid operator that is a modified conservative semi-Lagrangian discretization and generalizes those we have developed previously for nonconservative scalar linear hyperbolic problems. Numerical tests are performed for the inviscid Burgers and Buckley–Leverett equations. For many test problems, the solver converges in just a handful of iterations with a convergence rate independent of mesh resolution, including problems with (interacting) shocks and rarefactions.

97 MATHEMATICS AND COMPUTING↗

Extending HPF for advanced data parallel applications

The stated goal of High Performance Fortran (HPF) was to 'address the problems of writing data parallel programs where the distribution of data affects performance'. After examining the current version of the language we are led to the conclusion that HPF has not fully achieved this goal. While the basic distribution functions offered by the language - regular block, cyclic, and block cyclic distributions - can support regular numerical algorithms, advanced applications such as particle-in-cell codes or unstructured mesh solvers cannot be expressed adequately. We believe that this is a major weakness of HPF, significantly reducing its chances of becoming accepted in the numeric community. The paper discusses the data distribution and alignment issues in detail, points out some flaws in the basic language, and outlines possible future paths of development. Furthermore, we briefly deal with the issue of task parallelism and its integration with the data parallel paradigm of HPF.

Chapman, Barbara↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗