Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Effect of current on spectrum of breaking waves in water of finite depth

This paper presents an approximate method to compute the mean value, the mean square value and the spectrum of waves in water of finite depth taking into account the effect of wave breaking with or without the presence of current. It is assumed that there exists a linear and Gaussian ideal wave train whose spectrum is first obtained using the wave energy flux balance equation without considering wave breaking. The Miche wave breaking criterion for waves in finite water depth is used to limit the wave elevation and establish an expression for the breaking wave elevation in terms of the elevation and its second time derivative of the ideal waves. Simple expressions for the mean value, the mean square value and the spectrum are obtained. These results are applied to the case in which a deep water unidirectional wave train, propagating normally towards a straight shoreline over gently varying sea bottom of parallel and straight contours, encounters an adverse steady current whose velocity is assumed to be uniformly distributed with depth. Numerical results are obtained and presented in graphical form.

Tung, C. C.↗

Snow ALbedo eVOlution (SALVO) Campaign Spectral Albedo and Related Measurements from April - June, 2024 in Utqiagivk, AK

A field-portable spectroradiometer, referred to herein as an ‘ASD’, was used to make spatially-distributed spectral albedo (350 – 2500 nm) measurements on tundra and sea ice surfaces. The ASD detector is carried in a backpack and controlled via a computer mounted on the front of the operator (see Figure 1). The ASD measures the spectral irradiance from a fiber optic cable that is routed from the backpack to a custom, gooseneck cosine collector mounted on the end of a 1-m long boom (Grenfell and Perovich, 2008). The boom was held at hip height (approximately 1 m) and had an integrated bubble level for levelling. To make an albedo measurement, first the operator collect an incident (down-welling) irradiance, followed by a reflected (up-welling) measurement. The time between incident and reflected measurements was typically between 11 and 26 seconds (interquartile range). For each measurement, 10 spectra are averaged together. Albedo is calculated as the ratio of the reflected to incident measurement, which obviates the need for absolute radiometric calibration. Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1 m south of the line. While the ASD operator was making measurements, an assistant kept notes on the scan number associated with each measurement, the surface type (see below), and collected photos of each measurement (see companion oblique photos data archive). Measurements were made within 3 hours of solar noon.

ASD Spectroradiometer↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

A DSMC Surface Chemistry Model for Carbon-Based Ablators

A detailed molecular surface chemistry model for the DSMC (Direct Simulation Monte Carlo) method is proposed and implemented into the SPARTA (Stochastic PArallel Rarefied-gas Time-accurate Analyzer) DSMC solver. Molchanova et al. constructed a molecular model for surface recombination in DSMC that includes different surface processes (adsorption, desoprtion, Eley-Rideal and Langmuir-Hinshelwood). All surface processes can be divided into two groups: surface mechanisms, which involve only the particle adsorbed by the surface (desorption and Langmuir-Hinshelwood), and impact mechanisms, which also involve gas-phase particles (adsorption, Eley-Rideal). Using a similar approach, the 14-reaction kinetic model of oxygen-carbon interaction suggested by Zhlukhtov and Abe, as well as more recent models by Alba et al., Poovathinghal et al., and a new model developed in the scope of this work, are implemented in SPARTA. The computational results for the different oxidation models are compared with experimental results from Murray et al. (oxidation of a vitreous carbon surface due to a hyperthermal beam of O and O2), with a particular focus on fluxes, angular and Time-Of-Flight distributions of scattered particles.

oxidation↗

Control of multiple resonant power processors in a multi-source system

Analysis and test results show that phasor-regulated, Mapham-derived resonant inverters can be paralleled to provide standardizing interfaces for multiple sources on a utility-type, aerospace power distribution bus. The basic sources do not require matching in any way, and may have grossly different characteristics. Fully stable system architectures with multiple sources, parallel/redundant distribution buses, and a wide variety of loads can be easily constructed and controlled. The commands and parameters available for system control allow for tight tolerance bus voltage control, and absolute power-sharing control from the various sources over the full range of possible source and load variations. That level of control enables simplified load power processing hardware and the distribution of losses to optimally load the source thermal control system. Positive control of all system performance and allocation of losses are not required by all missions or vehicles, and overall vehicle considerations do not always require the loads on vehicle energy sources and thermal control systems to be balanced. In those cases, power system control can be simplified, and a hierarchical set of defaults can be substituted for computer-generated or supervisory input commands to allow for stable, fully autonomous system operation.

Mildice, James↗

Impact of Reordering on the LU Factorization Performance of Bordered Block-Diagonal Sparse Matrix

Power engineers rely on computer-based simulation tools to assess grid performance and ensure security. At the core of these tools are solvers for sparse linear equations. When transformed into a bordered block-diagonal (BBD) structure, part of the sparse linear equation solving can be parallelized. This work focuses on using the Schur-complement-based method for LU factorization on BBD matrices, specifically, Jacobian matrices from large-scale systems. Our findings show that the natural ordering method outperforms the default ordering method in computational performance for each block of the BBD matrix. This observation is validated using synthetic 25k-bus and 70k-bus cases, showing a speedup of up to 38% when using natural ordering without permutation. Additionally, the impact of the number of partitions is studied, and the result shows that computational performance improves with more, smaller partitions in the BBD matrices.

BBD matrix↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.5)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

High-performance mass storage system for workstations

Reduced Instruction Set Computer (RISC) workstations and Personnel Computers (PC) are very popular tools for office automation, command and control, scientific analysis, database management, and many other applications. However, when using Input/Output (I/O) intensive applications, the RISC workstations and PC's are often overburdened with the tasks of collecting, staging, storing, and distributing data. Also, by using standard high-performance peripherals and storage devices, the I/O function can still be a common bottleneck process. Therefore, the high-performance mass storage system, developed by Loral AeroSys' Independent Research and Development (IR&D) engineers, can offload a RISC workstation of I/O related functions and provide high-performance I/O functions and external interfaces. The high-performance mass storage system has the capabilities to ingest high-speed real-time data, perform signal or image processing, and stage, archive, and distribute the data. This mass storage system uses a hierarchical storage structure, thus reducing the total data storage cost, while maintaining high-I/O performance. The high-performance mass storage system is a network of low-cost parallel processors and storage devices. The nodes in the network have special I/O functions such as: SCSI controller, Ethernet controller, gateway controller, RS232 controller, IEEE488 controller, and digital/analog converter. The nodes are interconnected through high-speed direct memory access links to form a network. The topology of the network is easily reconfigurable to maximize system throughput for various applications. This high-performance mass storage system takes advantage of a 'busless' architecture for maximum expandability. The mass storage system consists of magnetic disks, a WORM optical disk jukebox, and an 8mm helical scan tape to form a hierarchical storage structure. Commonly used files are kept in the magnetic disk for fast retrieval. The optical disks are used as archive media, and the tapes are used as backup media. The storage system is managed by the IEEE mass storage reference model-based UniTree software package. UniTree software will keep track of all files in the system, will automatically migrate the lesser used files to archive media, and will stage the files when needed by the system. The user can access the files without knowledge of their physical location. The high-performance mass storage system developed by Loral AeroSys will significantly boost the system I/O performance and reduce the overall data storage cost. This storage system provides a highly flexible and cost-effective architecture for a variety of applications (e.g., realtime data acquisition with a signal and image processing requirement, long-term data archiving and distribution, and image analysis and enhancement).

Chiang, T.↗

3D modeling of deep borehole electromagnetic measurements with energized casing source for fracture mapping at the Utah Frontier Observatory for Research in Geothermal Energy

Here, we present a 3D numerical modelling analysis evaluating the deployment of a borehole electromagnetic measurement tool to detect and image a stimulated zone at the Utah Frontier Observatory for Research in Geothermal Energy geothermal site. As the depth to the geothermal reservoir is several kilometres and the size of the stimulated zone is limited to several 100 m, surface-based controlled-source electromagnetic measurements lack the sensitivity for detecting changes in electrical resistivity caused by the stimulation. To overcome the limitation, the study evaluates the feasibility of using a three-component borehole magnetic receiver system at the Frontier Observatory for Research in Geothermal Energy site. To provide sufficient currents inside and around the enhanced geothermal reservoir, we use an injection well as an energized casing source. To efficiently simulate energizing the injection well in a realistic 3D resistivity model, we introduce a novel modelling workflow that leverages the strengths of both 3D cylindrical-mesh-based electromagnetic modelling code and 3D tetrahedral-mesh-based electromagnetic modelling code. The former is particularly well-suited for modelling hollow cylindrical objects like casings, whereas the latter excels at representing more complex 3D geological structures. In this workflow, our initial step involves computing current densities along a vertical steel-cased well using a 3D cylindrical electromagnetic modelling code. Subsequently, we distribute a series of equivalent current sources along the well's trajectory within a complex 3D resistivity model. We then discretize this model using a tetrahedral mesh and simulate the borehole electromagnetic responses excited by the casing source using a 3D finite-element electromagnetic code. This multi-step approach enables us to simulate 3D casing source electromagnetic responses within a complex 3D resistivity model, without the need for explicit discretization of the well using an excessive number of fine cells. We discuss the applicability and limitations of this proposed workflow within an electromagnetic modelling scenario where an energized well is deviated, such as at the Frontier Observatory for Research in Geothermal Energy site. Using the workflow, we demonstrate that the combined use of the energized casing source and the borehole electromagnetic receiver system offer measurable magnetic field amplitudes and sensitivity to the deep localized stimulated zone. The measurements can also distinguish between parallel-fracture anisotropic reservoirs and isotropic cases, providing valuable insights into the fracture system of the stimulated zone. Besides the magnetic field measurements, vertical electric field measurements in the open well sections are also highly sensitive to the stimulated zone and can be used as additional data for detecting and imaging the target. We can also acquire additional multiple-source data by grounding the surface electrode at various locations and repeating borehole electromagnetic measurements. This approach can increase the number of monitoring data by several factors, providing a more comprehensive dataset for analysing the deep-localized stimulated zone. The numerical analysis indicates that it is feasible to use the combination of the energized casing and downhole electromagnetic measurements in monitoring localized stimulated zone at large depths.

58 GEOSCIENCES↗

Parallelization of Rocket Engine Simulator Software (PRESS)

We have outlined our work in the last half of the funding period. We have shown how a demo package for RESSAP using MPI can be done. However, we also mentioned the difficulties with the UNIX platform. We have reiterated some of the suggestions made during the presentation of the progress of the at Fourth Annual HBCU Conference. Although we have discussed, in some detail, how TURBDES/PUMPDES software can be run in parallel using MPI, at present, we are unable to experiment any further with either MPI or PVM. Due to X windows not being implemented, we are also not able to experiment further with XPVM, which it will be recalled, has a nice GUI interface. There are also some concerns, on our part, about MPI being an appropriate tool. The best thing about MPr is that it is public domain. Although and plenty of documentation exists for the intricacies of using MPI, little information is available on its actual implementations. Other than very typical, somewhat contrived examples, such as Jacobi algorithm for solving Laplace's equation, there are few examples which can readily be applied to real situations, such as in our case. In effect, the review of literature on both MPI and PVM, and there is a lot, indicate something similar to the enormous effort which was spent on LISP and LISP-like languages as tools for artificial intelligence research. During the development of a book on programming languages [12], when we searched the literature for very simple examples like taking averages, reading and writing records, multiplying matrices, etc., we could hardly find a any! Yet, so much was said and done on that topic in academic circles. It appears that we faced the same problem with MPI, where despite significant documentation, we could not find even a simple example which supports course-grain parallelism involving only a few processes. From the foregoing, it appears that a new direction may be required for more productive research during the extension period (10/19/98 - 10/18/99). At the least, the research would need to be done on Windows 95/Windows NT based platforms. Moreover, with the acquisition of Lahey Fortran package for PC platform, and the existing Borland C + + 5. 0, we can do work on C + + wrapper issues. We have carefully studied the blueprint for Space Transportation Propulsion Integrated Design Environment for the next 25 years [13] and found the inclusion of HBCUs in that effort encouraging. Especially in the long period for which a map is provided, there is no doubt that HBCUs will grow and become better equipped to do meaningful research. In the shorter period, as was suggested in our presentation at the HBCU conference, some key decisions regarding the aging Fortran based software for rocket propellants will need to be made. One important issue is whether or not object oriented languages such as C + + or Java should be used for distributed computing. Whether or not "distributed computing" is necessary for the existing software is yet another, larger, question to be tackled with.

Cezzar, Ruknet↗

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

Root-Raised Cosine Filter Implementation That Uses Canonical Signed Digits for High-Speed Digital Filter Applications

NASA Lewis Research Center's Space Communications Division has been investigating high-speed digital filters that can operate at a higher speed than those in current use for a digital modulator and demodulator (modem). Using the Canonical Signed Digits (CSD) number representation for filter coefficients is a very effective way to increase the filter's speed while reducing complexity in the digital filter hardware design. This approach is a good alternative to using an expensive parallel-processing design technique or custom, application-specific integrated circuits. Such integrated circuits may not be suitable for applications that require filter speeds faster than what application-specific integrated circuits digital signal processors can offer for a dedicated channel. When a communication channel is a dedicated, multiplication process--a costly, time-consuming process--it can be greatly simplified by a replacement of the filter coefficients with CSD numbers. A computer code written with the MATLAB software package runs the program and generates CSD-represented filter coefficients that are based on minimizing minimum mean square errors. Also, the Alta Group of Cadence's Signal Processing Workstation is used to simulate and analyze the CSD filter responses. The impulse response of the root-raised cosine filter that is used as a base model is defined. From this filter, a set of coefficients is sampled and stored in a file. For the all coefficients, the optimal CSD number for each coefficient is searched on the basis of the minimum-mean-square-errors criterion. Because the distribution of CSD numbers is not uniform, quantization errors tend to be bigger for coefficients greater than 1/2. To offset errors that occur in a region of coefficients between 1/2 to 1 and to better represent fractions with CSD numbers, an extra nonzero digit is allowed for any coefficients exceeding 1/2. This will greatly improve frequency response as well as intersymbol interference at the receiver. The frequency response of a set of collected CSD-represented filter coefficients was compared with the same filter that was conventionally implemented. Analyses show CSD-implemented filters perform as well as conventional filters. Comparison of eye diagrams and bit-error-rate curves between CSD filters and traditionally implemented filters are almost indistinguishable. However, filter complexity was reduced from almost 3.5 to 1 for CSD filters. Complete computer simulation results are available. In the near future, work will focus on building actual working digital filter hardware in a field programmable gate array (FPGA).

Kim, Heechul↗

Development of a Detailed Surface Chemistry Framework in DSMC

A generalized finite-rate surface chemistry framework incorporating a comprehensive list of reaction mechanisms is developed and implemented into the Direct Simulation Monte Carlo (DSMC) solver SPARTA (Stochastic PArallel Rarefied-gas Time-accurate Analyzer). The various mechanisms include adsorption, desorption, Eley-Rideal (ER), and several types of Langmuir-Hinshelwood (LH) mechanisms. The approach is to stochastically model the various competing reactions occurring on a set of active sites. Both gas-surface (e.g., adsorption, ER) and pure-surface (e.g., desorption) reaction mechanisms are incorporated, and the framework also includes catalytic or surface altering mechanisms involving the participation of the bulk-phase species (e.g., bulk carbon atoms). Marschall and MacLean developed a general formulation in which multiple phases and surface sites are used and a similar convention is adopted in the current work. Expressions for the microscopic parameters of reaction probabilities (for gas-surface reactions) and frequencies (for pure-surface reactions) that are required for DSMC are derived from the surface properties and macroscopic parameters such as rate constants, sticking coefficients, etc. The energy and angular distributions of the products are specified according to the reaction type and input parameters. This framework also presents physically consistent procedures to accurately compute the reaction probabilities and frequencies in the case of multiple reactions. The result is a modeling tool with a wide variety of surface reactions characterized via user-specified reaction rate constants, surface properties and parameters.

Surface Chemistry↗

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

A Full-Stack Exploration of Language-Based Parallelism in Fortran 2023

This poster explores native parallel features in Fortran 2023 through the lens of supporting applications with libraries, compilers, and parallel runtimes. The language revision informally named Fortran 2008 introduced parallelism in the form of Single Program Multiple Data (SPMD) execution with two broad feature sets: (1) loop-level parallelism via do concurrent and (2) a Partitioned Global Address Space (PGAS) comprised of distributed “coarray” data structures. Fortran’s native parallelism has demonstrated high performance [1] and reduced the burden of inserting what sometimes amounts to more directives than code. Several compilers support both feature sets, typically by translating do concurrent into serial do loops annotated by parallel directives and by translating SPMD/PGAS features into direct calls to a communication library. Our research focuses primarily on two questions: (1) can the compiler’s parallel runtime library be developed in the language being compiled (Fortran) and (2) can we define an interface to the runtime that liberates compilers from being hardwired to one runtime and vice versa. We are answering these questions by developing the Parallel Runtime Interface for Fortran (PRIF) [2] and the Co-Array Fortran Framework of Efficient Interfaces to Network Environments (Caffeine) [3]. Caffeine is initially targeting adoption by LLVM Flang, a new open-source Fortran compiler developed by a broad community in industry, academia, and government labs. We are also exploring the use of these features in Inference-Engine, a deep learning library designed to facilitate neural network training and inference for high-performance computing applications written in modern Fortran.

Rasmussen, Katherine↗

0D, 1D, 2D, and 3D simulations of an idealized coaxial impedance-matched Marx generator

We have conducted 0D, 1D, 2D, and 3D simulations of an idealized coaxial impedance-matched Marx generator (IMG) []. The 0D calculations were conducted with a four-element circuit model; the 1D, 2D, and 3D calculations were conducted with highly resolved, fully electromagnetic representations. The IMG consists of 30 stages distributed axially and connected electrically in series. Each stage is powered by two bricks separated by 180° and connected electrically in parallel. Each brick comprises two opposite-polarity capacitors in series with a single switch. The bricks drive an internal impedance-matched coaxial transmission line terminated by a resistive load. The simulations neglect effects due to the switch-triggering circuit, the capacitor-charging circuit, external conducting boundaries, and reactive components of the load. We find dimensionality does not significantly affect the electrical power delivered by the IMG to its load: peak load powers estimated by the 0D, 1D, 2D, and 3D simulations agree to within 1%. The 3D calculations demonstrate that electromagnetic power radiated by the bricks, and axial gaps between stages, reduces the peak load power by less than ∼ 1 % . Each simulation assumes the load impedance is 34% above that at which the load power is maximized. Operating an IMG with such an overmatched load offers several advantages while decreasing the peak load power by only 2%. The 0D, 1D, 2D, and 3D models outlined herein could be adapted to assess computationally competing IMG designs, and conduct a variety of numerical IMG experiments, an IMG is constructed. Published by the American Physical Society 2024

43 PARTICLE ACCELERATORS↗

Performance of BLAS 3, FFTs and NAS Parallel Benchmarks on Cray T3D

Recently, a Cray T3D Emulator has been made available on the Cray Y-MP and C90 computers. The Pittsburgh Supercomputer Center has acquired a CRAY T3D system and many other centers like Jet Propulsion Laboratory (JPL) will have it by the end of 1994. The Cray T3D system is the firstphase system in Cray Research, Inc.'s (CRI) three-phase massively parallel processing (MPP) program. This system features a heterogeneous architecture that closely couples DEC's ALPHA microprocessors and CRI's parallel-vector technology, i.e. the Cray Y-MP and Cray C90. The Cray T3D Emulator will give prospective users a valuable experience in developing high performance applications on the MPP system. This emulator runs programs written in CRI's MPP Fortran programming model (data sharing and work sharing) or Parallel Virtual Machine (PVM) programming model. It will help the users to study data layout, data locality, and data reference patterns thereby providing feedback which will enable one to write more efficient parallel codes. An overview of the Cray T3D hardware, software, and three of its available programming models is presented.The Cray Fortran Programming Model comprising (a) Data Sharing, (b) Worksharing and (c) Message Passing, will be discussed with examples. We have also implemented distributed BLAS 3 (matrix-matrix multiplication) in data parallel model (using only CSHIFT); worksharing model using block distribution and collapsed distribution; and message passing model using PVM. We have also implemented 2D and 3D FFTs for radix-2 using PVM. The performance of NAS Parallel 'Benchmarks (NPB) on CRAY T3D will be compared with other highly parallel systems such as CM-5, Paragon, C90 etc.

Saini, Subhash↗