Search NASA⌕ Search

SEARCH · Search NASA

Results for “Distributed computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

The Turbulent Flow in Diffusers of Small Divergence Angle

The turbulent flow in a conical diffuser represents the type of turbulent boundary layer with positive longitudinal pressure gradient. In contrast to the boundary layer problem, however, it is not necessary that the pressure distribution along the limits of the boundary layer(along the axis of the diffuser) be given, since this distribution can be obtained from the computation. This circumstance, together with the greater simplicity of the problem as a whole, provides a useful basis for the study of the extension of the results of semiempirical theories to the case of motion with a positive pressure gradient. In the first part of the paper,formulas are derived for the computation of the velocity and.pressure distributions in the turbulent flow along, and at right angles to, the axis of a diffuser of small cone angle. The problem is solved.

Gourzhienko, G. A.↗

An Offload NIC for NASA, NLR, and Grid Computing

This work addresses distributed data management and access dynamically configurable high-speed access to data distributed and shared over wide-area high-speed network environments. An offload engine NIC (network interface card) is proposed that scales at nX10-Gbps increments through 100-Gbps full duplex. The Globus de facto standard was used in projects requiring secure, robust, high-speed bulk data transport. Novel extension mechanisms were derived that will combine these technologies for use by GridFTP, bandwidth management resources, and host CPU (central processing unit) acceleration. The result will be wire-rate encrypted Globus grid data transactions through offload for splintering, encryption, and compression. As the need for greater network bandwidth increases, there is an inherent need for faster CPUs. The best way to accelerate CPUs is through a network acceleration engine. Grid computing data transfers for the Globus tool set did not have wire-rate encryption or compression. Existing technology cannot keep pace with the greater bandwidths of backplane and network connections. Present offload engines with ports to Ethernet are 32 to 40 Gbps f-d at best. The best of ultra-high-speed offload engines use expensive ASICs (application specific integrated circuits) or NPUs (network processing units). The present state of the art also includes bonding and the use of multiple NICs that are also in the planning stages for future portability to ASICs and software to accommodate data rates at 100 Gbps. The remaining industry solutions are for carrier-grade equipment manufacturers, with costly line cards having multiples of 10-Gbps ports, or 100-Gbps ports such as CFP modules that interface to costly ASICs and related circuitry. All of the existing solutions vary in configuration based on requirements of the host, motherboard, or carriergrade equipment. The purpose of the innovation is to eliminate data bottlenecks within cluster, grid, and cloud computing systems, and to add several more capabilities while reducing space consumption and cost. Provisions were designed for interoperability with systems used in the NASA HEC (High-End Computing) program. The new acceleration engine consists of state-ofthe- art FPGA (field-programmable gate array) core IP, C, and Verilog code; novel communication protocol; and extensions to the Globus structure. The engine provides the functions of network acceleration, encryption, compression, packet-ordering, and security added to Globus grid or for cloud data transfer. This system is scalable in nX10-Gbps increments through 100-Gbps f-d. It can be interfaced to industry-standard system-side or network-side devices or core IP in increments of 10 GigE, scaling to provide IEEE 40/100 GigE compliance.

Awrach, James↗

Turbulent Navier-Stokes Flow Analysis of an Advanced Semispan Diamond-Wing Model in Tunnel and Free Air at High-Lift Conditions

Turbulent Navier-Stokes computational results are presented for an advanced diamond wing semispan model at low-speed, high-lift conditions. The numerical results are obtained in support of a wind-tunnel test that was conducted in the National Transonic Facility at the NASA Langley Research Center. The model incorporated a generic fuselage and was mounted on the tunnel sidewall using a constant-width non-metric standoff. The computations were performed at to a nominal approach and landing flow conditions.The computed high-lift flow characteristics for the model in both the tunnel and in free-air environment are presented. The computed wing pressure distributions agreed well with the measured data and they both indicated a small effect due to the tunnel wall interference effects. However, the wall interference effects were found to be relatively more pronounced in the measured and the computed lift, drag and pitching moment. Although the magnitudes of the computed forces and moment were slightly off compared to the measured data, the increments due the wall interference effects were predicted reasonably well. The numerical results are also presented on the combined effects of the tunnel sidewall boundary layer and the standoff geometry on the fuselage forebody pressure distributions and the resulting impact on the configuration longitudinal aerodynamic characteristics.

Ghaffari, Farhad↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Tensor network simulations of quasi-GPDs in the massive Schwinger model

Generalized parton distribution functions (GPDs) are off-diagonal light-cone matrix elements that encode the internal structure of hadrons in terms of quark and gluon degrees of freedom. In this work, we present the first nonperturbative study of quasi-GPDs in the massive Schwinger model, quantum electrodynamics in 1+1 dimensions (QED 2 ), within the Hamiltonian formulation of lattice field theory. Quasidistributions are spatial correlation functions of boosted states, which approach the relevant light-cone distributions in the luminal limit. Using tensor networks, we prepare the first excited state in the strongly coupled regime and boost it to close to the light-cone on lattices of up to 400 lattice sites. We compute both quasiparton distribution functions and, for the first time, quasi-GPDs, and study their convergence for increasingly boosted states. In addition, we perform analytic calculations of GPDs in the two-particle Fock-space approximation and in the Reggeized limit, providing qualitative benchmarks for the tensor network results. Our analysis establishes computational benchmarks for accessing partonic observables in low-dimensional gauge theories, offering a starting point for future extensions to higher dimensions, non-Abelian theories, and quantum simulations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Research in Parallel Algorithms and Software for Computational Aerosciences

Phase I is complete for the development of a Computational Fluid Dynamics parallel code with automatic grid generation and adaptation for the Euler analysis of flow over complex geometries. SPLITFLOW, an unstructured Cartesian grid code developed at Lockheed Martin Tactical Aircraft Systems, has been modified for a distributed memory/massively parallel computing environment. The parallel code is operational on an SGI network, Cray J90 and C90 vector machines, SGI Power Challenge, and Cray T3D and IBM SP2 massively parallel machines. Parallel Virtual Machine (PVM) is the message passing protocol for portability to various architectures. A domain decomposition technique was developed which enforces dynamic load balancing to improve solution speed and memory requirements. A host/node algorithm distributes the tasks. The solver parallelizes very well, and scales with the number of processors. Partially parallelized and non-parallelized tasks consume most of the wall clock time in a very fine grain environment. Timing comparisons on a Cray C90 demonstrate that Parallel SPLITFLOW runs 2.4 times faster on 8 processors than its non-parallel counterpart autotasked over 8 processors.

Domel, Neal D.↗

Research in Parallel Algorithms and Software for Computational Aerosciences

Phase 1 is complete for the development of a computational fluid dynamics CFD) parallel code with automatic grid generation and adaptation for the Euler analysis of flow over complex geometries. SPLITFLOW, an unstructured Cartesian grid code developed at Lockheed Martin Tactical Aircraft Systems, has been modified for a distributed memory/massively parallel computing environment. The parallel code is operational on an SGI network, Cray J90 and C90 vector machines, SGI Power Challenge, and Cray T3D and IBM SP2 massively parallel machines. Parallel Virtual Machine (PVM) is the message passing protocol for portability to various architectures. A domain decomposition technique was developed which enforces dynamic load balancing to improve solution speed and memory requirements. A host/node algorithm distributes the tasks. The solver parallelizes very well, and scales with the number of processors. Partially parallelized and non-parallelized tasks consume most of the wall clock time in a very fine grain environment. Timing comparisons on a Cray C90 demonstrate that Parallel SPLITFLOW runs 2.4 times faster on 8 processors than its non-parallel counterpart autotasked over 8 processors.

Domel, Neal D.↗

MixPI: Mixed-time slicing path integral software for quantized molecular dynamics simulations

We introduce the MixPI software to implement path integral molecular dynamics (PIMD) simulations for the study of condensed phase systems where nuclear quantum effects (NQEs) are important. In contrast to existing PIMD simulation software, MixPI enables the implementation of mixed quantum–classical path integral simulations where only a subset of system degrees of freedom (dofs) are treated quantum mechanically in an extended phase space while the remaining dofs are described classically. We expect this software to be particularly useful for simulations of electron and proton transfer in condensed phase systems, as well as for the study of biological and material systems where only a handful of dofs contribute significantly to the observed NQEs. We demonstrate the use of MixPI in two different systems. The first is a simple water model where we implement a set of mixed quantum–classical simulations to compute average energy and radial distribution functions. We use these simulations to benchmark the effectiveness of MixPI and to demonstrate how it enables systematic investigation into the origin of observed NQEs. We then compute radial distribution functions for a system where MixPI is essential: a solvated metal (M 2+ ) cation described using an explicit quantized electron localized on an M 3+ ion in water.

chemical physics↗

Collecting address traces from parallel computers

Trace driven simulation is a well-established method of performance analysis for single processor computer systems. However, efficient and accurate memory address tracing for parallel computer systems is not well understood. In this paper we present a critical survey of recently implemented approaches to address tracing and highlight the issues specific to collection of traces for both shared and distributed memory parallel computers. These issues include potential distortion of the relative ordering of events by the address tracing activity, realistic interleaving of addresses generated by multiple processors, and I/O and storage problems associated with collecting traces for large parallel systems. The strengths and weaknesses of the parallel tracing approaches are described.

Stunkel, Craig B.↗

Calculating Cumulative Binomial-Distribution Probabilities

Cumulative-binomial computer program, CUMBIN, one of set of three programs, calculates cumulative binomial probability distributions for arbitrary inputs. CUMBIN, NEWTONP (NPO-17556), and CROSSER (NPO-17557), used independently of one another. Reliabilities and availabilities of k-out-of-n systems analyzed. Used by statisticians and users of statistical procedures, test planners, designers, and numerical analysts. Used for calculations of reliability and availability. Program written in C.

Scheuer, Ernest M.↗

SCALE 6.3 Validation: Radiation Shielding

Safe and reliable use of scientific and engineering computer codes requires validation for the types of applications in which they will be used. An example in the nuclear reactor engineering and licensing field is radiation transport employed in shielding analyses. The validity of computer codes for shielding applications is demonstrated in this report for SCALE version 6.3.0. Representative benchmarks corresponding to shielding analyses are selected for the validation study. Typical measurement results analyzed from these benchmarks include neutron fluxes, detector count rates, detector energy response functions, neutron and gamma dose rates, neutron activation rates and activities, neutron leakage fluxes, and skyshine dose rates. Thousands of points of comparison between measurement and calculation are presented in this work. Other than rare outliers typically explained by either a lack of information or large uncertainties in the experiment conditions, material, or dimensions, the Monaco with Automated Variance Reduction using Importance Calculations (MAVRIC) radiation transport computer code with built-in variance reduction methods distributed with the SCALE computer code system agrees well with the measurement results. In selected benchmarks, MAVRIC is also compared to Monte Carlo N- Particle® (MCNP® ) 1 calculations. Both computer codes generally agree well within the estimated uncertainties. With the release of SCALE 6.3.0, Shift was integrated as an alternative transport solver in MAVRIC, denoted MAVRIC-Shift. Although the traditional MAVRIC using Monaco was used primarily in this validation study, many results have also been generated using MAVRIC-Shift. Agreement between MAVRIC-Monaco and MAVRIC-Shift is generally very good. The benchmarks presented in this report were obtained from reliable sources such as the International Criticality Safety Benchmark Evaluation Project Handbook, the Shielding Integral Benchmark Archive & Database, and other shielding validation work found in the literature. Additional datapoints and benchmarks will be added to future versions of this report to expand the shielding validation suite.

61 RADIATION PROTECTION AND DOSIMETRY↗

Supercritical flow about a thick circular-arc airfoil

The supercritical flow about a biconvex circular-arc airfoil is being thoroughly documented at Ames Research Center in order to provide experimental test cases suitable for guiding and evaluating current and future computer codes. The effects of angle of attack, effects of leading and trailing-edge splitter plates, additional unsteady pressure fluctuation (buffeting) measurements and glow-field shadowgraphs, and application of an oil-film technique to display separated-wake streamlines were studied. Computed and measured pressure distributions for steady and unsteady flows, using a recent computer code representative of current methodology, are compared. It was found that the numerical solutions are often fundamentally incorrect in that only strong (shock-polar terminology) shocks are captured, whereas experimentally, both strong and weak shock waves appear.

Mcdevitt, J. B.↗

An analysis of the viscous flow through a compact radial turbine by the average passage approach

A steady, three-dimensional viscous average passage computer code is used to analyze the flow through a compact radial turbine rotor. The code models the flow as spatially periodic from blade passage to blade passage. Results from the code using varying computational models are compared with each other and with experimental data. These results include blade surface velocities and pressures, exit vorticity and entropy contour plots, shroud pressures, and spanwise exit total temperature, total pressure, and swirl distributions. The three computational models used are inviscid, viscous with no blade clearance, and viscous with blade clearance. It is found that modeling viscous effects improves correlation with experimental data, while modeling hub and tip clearances further improves some comparisons. Experimental results such as a local maximum of exit swirl, reduced exit total pressures at the walls, and exit total temperature magnitudes are explained by interpretation of the flow physics and computed secondary flows. Trends in the computed blade loading diagrams are similarly explained.

Heidmann, James D.↗

An analysis of the viscous flow through a compact radial turbine by the average passage approach

A steady, three-dimensional viscous average passage computer code is used to analyze the flow through a compact radial turbine rotor. The code models the flow as spatially periodic from blade passage to blade passage. Results from the code using varying computational models are compared with each other and with experimental data. These results include blade surface velocities and pressures, exit vorticity and entropy contour plots, shroud pressures, and spanwise exit total temperature, total pressure, and swirl distributions. The three computational models used are inviscid, viscous with no blade clearance, and viscous with blade clearance. It is found that modeling viscous effects improves correlation with experimental data, while modeling hub and tip clearances further improves some comparisons. Experimental results such as a local maximum of exit swirl, reduced exit total pressures at the walls, and exit total temperature magnitudes are explained by interpretation of the flow physics and computed secondary flows. Trends in the computed blade loading diagrams are similarly explained.

Heidmann, James D.↗