Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

Applications considerations in the system design of highly concurrent multiprocessors

A flow model processor approach to parallel processing is described, using very-high-performance individual processors, high-speed circuit switched interconnection networks, and a high-speed synchronization capability to minimize the effect of the inherently serial portions of applications on performance. Design studies related to the determination of the number of processors, the memory organization, and the structure of the networks used to interconnect the processor and memory resources are discussed. Simulations indicate that applications centered on the large shared data memory should be able to sustain over 500 million floating point operations per second.

Lundstrom, Stephen F.↗

A Comparison of Three Programming Models for Adaptive Applications

We study the performance and programming effort for two major classes of adaptive applications under three leading parallel programming models. We find that all three models can achieve scalable performance on the state-of-the-art multiprocessor machines. The basic parallel algorithms needed for different programming models to deliver their best performance are similar, but the implementations differ greatly, far beyond the fact of using explicit messages versus implicit loads/stores. Compared with MPI and SHMEM, CC-SAS (cache-coherent shared address space) provides substantial ease of programming at the conceptual and program orchestration level, which often leads to the performance gain. However it may also suffer from the poor spatial locality of physically distributed shared data on large number of processors. Our CC-SAS implementation of the PARMETIS partitioner itself runs faster than in the other two programming models, and generates more balanced result for our application.

Shan, Hong-Zhang↗

Pre-Stall Behavior of a Transonic Axial Compressor Stage via Time-Accurate Numerical Simulation

CFD calculations using high-performance parallel computing were conducted to simulate the pre-stall flow of a transonic compressor stage, NASA compressor Stage 35. The simulations were run with a full-annulus grid that models the 3D, viscous, unsteady blade row interaction without the need for an artificial inlet distortion to induce stall. The simulation demonstrates the development of the rotating stall from the growth of instabilities. Pressure-rise performance and pressure traces are compared with published experimental data before the study of flow evolution prior to the rotating stall. Spatial FFT analysis of the flow indicates a rotating long-length disturbance of one rotor circumference, which is followed by a spike-type breakdown. The analysis also links the long-length wave disturbance with the initiation of the spike inception. The spike instabilities occur when the trajectory of the tip clearance flow becomes perpendicular to the axial direction. When approaching stall, the passage shock changes from a single oblique shock to a dual-shock, which distorts the perpendicular trajectory of the tip clearance vortex but shows no evidence of flow separation that may contribute to stall.

Chen, Jen-Ping↗

An On Board Processor (OBP) for OAO C

A stored program computer and its application on OAO is considered. The parallel computer has a memory capacity of 16,384 words of 18 bits each, one central processor unit, two 4096 word memory units, and one input/output unit. The I/O has no direct data connection with the CPU so that all data flow between these two units must pass through memory by way of the memory data bus. The primary functions of the onboard computer are auxiliary command storage, spacecraft monitoring and malfunction reporting, data compression and status summary, and possible performance of emergency corrective action.

Hartenstein, R. G.↗

Molecular Dynamics Simulations of Silicon Carbide, Boron Nitride and Silicon for Ceramic Matrix Composite Applications

A comprehensive computational molecular dynamics study is presented for crystalline α-SiC (6H, 4H, and 2H SiC), β-SiC (3C SiC), layered boron nitride, amorphous boron nitride and silicon, the constituent materials for high-temperature SiC/SiC compositions. Large-scale Atomic/Molecular Parallel Simulator software package was used. The Tersoff Potential force field was utilized to evaluate their mechanical characteristics of most of the materials, and the Reax force field was used to model silicon when the Tersoff Potential did not provide accurate results. Their mechanical behaviors were evaluated at a strain rate of 10(exp 7)/s and the results agree with the experimental data in the literature. The results are foundational for linking constituent behavior to composite performance, particularly when test data is unavailable or suspect.

Aluko, Olanrewaju↗

Modeling of three-dimensional mixing and reacting ducted flows

A computer code based on a finite-element solution algorithm is developed to perform an analytical investigation on the turbulent mixing and reaction of hydrogen jets injected from multiple orifices transverse and parallel to a supersonic airstream. A laser optical cavity flow field was also analyzed to demonstrate the generality of the proposed model. Computational results provide a three-dimensional description of velocity, temperature, and species-concentration fields downstream of injection. Major conclusions are that the analysis has immediate utility in evaluating the mixing effectiveness of transverse H2 injection data since it has been tested in its ability to model this type of data and that turbulent mixing length theory, constant effective Prandtl number, and a Lewis number of unity provide reasonable agreement with transverse H2 injection data downstream of the near-injection region. Efforts are presently being directed toward using the code in modeling laser and scramjet data for a wide range of flow conditions.

Zelazny, S. W.↗

Fault-tolerant multichannel demultiplexer subsystems

Fault tolerance in future processing and switching communication satellites is addressed by showing new methods for detecting hardware failures in the first major subsystem, the multichannel demultiplexer. An efficient method for demultiplexing frequency slotted channels uses multirate filter banks which contain fast Fourier transform processing. All numerical processing is performed at a lower rate commensurate with the small bandwidth of each bandbase channel. The integrity of the demultiplexing operations is protected by using real number convolutional codes to compute comparable parity values which detect errors at the data sample level. High rate, systematic convolutional codes produce parity values at a much reduced rate, and protection is achieved by generating parity values in two ways and comparing them. Parity values corresponding to each output channel are generated in parallel by a subsystem, operating even slower and in parallel with the demultiplexer that is virtually identical to the original structure. These parity calculations may be time shared with the same processing resources because they are so similar.

Redinbo, Robert↗

Recent Advances in Photonic Devices for Optical Computing and the Role of Nonlinear Optics-Part II

The twentieth century has been the era of semiconductor materials and electronic technology while this millennium is expected to be the age of photonic materials and all-optical technology. Optical technology has led to countless optical devices that have become indispensable in our daily lives in storage area networks, parallel processing, optical switches, all-optical data networks, holographic storage devices, and biometric devices at airports. This chapters intends to bring some awareness to the state-of-the-art of optical technologies, which have potential for optical computing and demonstrate the role of nonlinear optics in many of these components. Our intent, in this Chapter, is to present an overview of the current status of optical computing, and a brief evaluation of the recent advances and performance of the following key components necessary to build an optical computing system: all-optical logic gates, adders, optical processors, optical storage, holographic storage, optical interconnects, spatial light modulators and optical materials.

Abdeldayem, Hossin↗

Simulation of Sweep-Jet Flow Control, Single Jet and Full Vertical Tail

This work is a simulation technology demonstrator, of sweep jets used to suppress boundary layer separation and increase maximum achievable load coefficients. A sweep jet is a discrete Coanda jet that oscillates in the plane parallel to an aerodynamic surface. It injects mass and momentum in the approximate stream wise direction. It also generate turbulent eddies at the oscillation frequency, which are typically large relative to boundary layer turbulence, and which augmenting mixing across the boundary layer to attack flow separation. Simulations of a fluidic oscillator, the sweep jet emerging from the oscillator, and the suppression of boundary layer separation by an array of sweep jets are performed. Simulation results are compared to data from a dedicated CFD validation experiment of a single oscillator and its sweep jet, and from a study of a full-scale Boeing 757 vertical tail augmented with an array of sweep jets.2, 20 A critical step in the work is the development of realistic time-dependent sweep-jet in flow boundary conditions, derived from the results of the single-oscillator simulations, which create the sweep jets in the full-tail simulations. Simulations were performed using the Over flow CFD solver, with high-order spatial discretization and a range of turbulence modeling. Good results were obtained for all flows simulated, when suitable turbulence modeling was used.

flow control↗

High-Rate Delay Tolerant Networking (HDTN) User Guide Version 1.0

Delay Tolerant Networking (DTN) has been identified as a key technology to enable and facilitate the development and growth of future space networks. Classically, space communications networks are collections of disparate links that are manually managed either point-to-point or use space relays. The accelerating accessibility of space enables a new scaling of space nodes, yet both the manual management of configurations and scheduling and the lack of structure connecting links precisely prohibit scaling. This challenge gives rise to newer and larger classes of communications needs that are met by DTN, which must overcome the disconnection, disruption, latency, and mobility featured in space communications systems. DTN joins the underlying links as an overlay, and can be made to communicate over any protocol stack. The core actions of DTN are store, carry, and forward, where data are stored instead of dropped if there is no immediately available outduct. It does this by taking the DTN unit of data, bundles, and providing necessary layers to adapt these bundles to the underlying transport protocols of choice; these are called convergence layers. DTN's Bundle Protocol (BP) can then be used on top of terrestrial protocol stacks, such as TCP/IP, as well as protocols for space, such as LTP/AOS, all in the same network. For emphasis it is noted that bundles can be of essentially any size, and hence this convergence to lower layers of choice is necessary. Existing DTN implementations have operated in constrained environments with limited resources, resulting in low data speeds. However, as various technologies have advanced, data transfer rates and efficiency have advanced, which has pushed the need for a DTN implementation for ground systems and for spacecraft that is performance-oriented in order to not impose an unnecessary bottleneck. High-rate Delay Tolerant Networking (HDTN) takes advantage of modern hardware platforms to substantially reduce latency and improve throughput compared to today’s DTN operations. The HDTN implementation maintains interoperability with existing deployments of DTN that conform to IETF RFCs 4838, 5050, and 9171. At the same time, HDTN defines a new data format better suited to higher-rate operation. It defines and adopts a massively parallel pipelined and message-oriented architecture, allowing the system to scale gracefully as its resources increase. HDTN’s architecture also supports hooks to replace various processing pipeline elements with specialized hardware accelerators. This offers improved Size, Weight, and Power (SWaP) characteristics while reducing development complexity and cost.

Delay Tolerant Networking↗

High-Rate Delay Tolerant Networking (HDTN) User Guide Version 1.3.0

Delay Tolerant Networking (DTN) has been identified as a key technology to enable and facilitate the development and growth of future space networks. Classically, space communications networks are collections of disparate links that are manually managed either point-to-point or use space relays. The accelerating accessibility of space enables a new scaling of space nodes, yet both the manual management of configurations and scheduling and the lack of structure connecting links precisely prohibit scaling. This challenge gives rise to newer and larger classes of communications needs that are met by DTN, which must overcome the disconnection, disruption, latency, and mobility featured in space communications systems. DTN joins the underlying links as an overlay, and can be made to communicate over any protocol stack. The core actions of DTN are store, carry, and forward, where data are stored instead of dropped if there is no immediately available outduct. It does this by taking the DTN unit of data, bundles, and providing necessary layers to adapt these bundles to the underlying transport protocols of choice; these are called convergence layers. DTN's Bundle Protocol (BP) can then be used on top of terrestrial protocol stacks, such as TCP/IP, as well as protocols for space, such as LTP/AOS, all in the same network. For emphasis it is noted that bundles can be of essentially any size, and hence this convergence to lower layers of choice is necessary. Existing DTN implementations have operated in constrained environments with limited resources, resulting in low data speeds. However, as various technologies have advanced, data transfer rates and efficiency have advanced, which has pushed the need for a DTN implementation for ground systems and for spacecraft that is performance-oriented in order to not impose an unnecessary bottleneck. High-rate Delay Tolerant Networking (HDTN) takes advantage of modern hardware platforms to substantially reduce latency and improve throughput compared to today’s DTN operations. The HDTN implementation maintains interoperability with existing deployments of DTN that conform to IETF RFCs 4838, 5050, and 9171. At the same time, HDTN defines a new data format better suited to higher-rate operation. It defines and adopts a massively parallel pipelined and message-oriented architecture, allowing the system to scale gracefully as its resources increase. HDTN’s architecture also supports hooks to replace various processing pipeline elements with specialized hardware accelerators. This offers improved Size, Weight, and Power (SWaP) characteristics while reducing development complexity and cost.

Delay Tolerant Networking↗

Simulating energetic ions and enhanced fusion rates from ion-cyclotron resonance heating with a full-wave/Fokker–Planck model

Reproducing fast-ion enhanced fusion rates from ion-cyclotron resonance heating (ICRH) in tokamaks requires the self-consistent coupling of a full-wave solver and a Fokker–Planck solver, which evolves multiple simultaneously resonant ion species. We introduce a new self-consistent model that iterates the TORIC full-wave solver with the CQL3D Fokker–Planck solver using the integrated plasma simulator (IPS). This model evolves the bounce-averaged ion distribution functions in both parallel and perpendicular velocity-space with a quasilinear radio frequency (RF) diffusion operator valid in the ion finite Larmor radius (FLR) limit and the RF electric fields with the resultant non-Maxwellian FLR dielectric tensor. This produces non-Maxwellian ICRH simulations that are fully self-consistent, fast, and interoperable with integrated modeling frameworks, such as TRANSP/GACODE/IPS-FASTRAN. We demonstrate our model's capabilities by validating it against experimental data in Alcator C-Mod. We then perform the first RF heating simulations of SPARC using self-consistent non-Maxwellian ion distributions to investigate the potential to enhance fusion rates using ion cyclotron resonance heating generated fast ions.

Physics↗

Process Design and Techno-Economic Analysis of the Modular Staged Pressurized Oxy-Combustion (SPOC) Power Plant for Biomass

This work describes the process design and techno-economic analysis (TEA) of the modular SPOC power plant for biomass firing and coal-biomass co-firing. Two Rankine cycles were considered: a supercritical steam cycle (242 bar, 593°C, 593°C) with 550 MWe net output and a subcritical cycle (166 bar, 566°C, 566°C) with 200 MWe net output. For both cases, 95% carbon capture was modeled, and hybrid poplar biomass was chosen to generate carbon-negative power. In addition, the supercritical 500 MWe case included a 25% biomass co-firing (carbon neutral) case. For both cycles, a 100% Powder River Basin coal firing case was used for comparison purposes. In the SPOC process, oxygen is produced via a cryogenic air separation unit (ASU) and the heat generated from the compression of air is integrated into the steam cycle and utilized for boiler feed water pre-heating. Unique to the SPOC process, the boilers are pressurized and arranged in a series-parallel configuration, with minimized flue gas recirculation. The flue gas is cooled and scrubbed in the direct-contact cooler (DCC) column, and the moisture in the flue gas is condensed, leaving the bottom of the DCC at a sufficiently high temperature such that it can be used for boiler feed water pre-heating, improving plant thermal efficiency. Following drying and purification, CO2 in the flue gas is at the purity required for storage or utilization. The performance data were obtained from process modelling via Aspen Plus®. The stream data from Aspen Plus® were used as an input for the AACE Class 5 cost study. Ultimately, the capital costs, Levelized Cost of Electricity (LCOE), and cost of CO2 captured and avoided were obtained. The HHV efficiency of the carbon negative 550 MWe supercritical SPOC case (34.8%) was clearly above those reported by NETL for the BECCS baseline cases of supercritical pulverized coal with capture (B12B, 31.5%) and the 49% biomass co-firing case with capture (PA3, 29.2%). The HHV efficiency of the carbon-negative subcritical plant is also higher than the subcritical baseline PC plant with capture (case B11B.95) presented by NETL (32% vs 29.7%). The LCOE for the SPOC 100% biomass case was similar to the LCOE for the BECCS 49% biomass with carbon capture case ($147/MWh), and the SPOC carbon neutral case LCOE was lower ($110/MWh) than the cost for the NETL baseline SC coal firing case with 90% carbon capture ($114/MWh).

Magalhaes, Duarte↗

Knowledge-based vision for space station object motion detection, recognition, and tracking

Computer vision, especially color image analysis and understanding, has much to offer in the area of the automation of Space Station tasks such as construction, satellite servicing, rendezvous and proximity operations, inspection, experiment monitoring, data management and training. Knowledge-based techniques improve the performance of vision algorithms for unstructured environments because of their ability to deal with imprecise a priori information or inaccurately estimated feature data and still produce useful results. Conventional techniques using statistical and purely model-based approaches lack flexibility in dealing with the variabilities anticipated in the unstructured viewing environment of space. Algorithms developed under NASA sponsorship for Space Station applications to demonstrate the value of a hypothesized architecture for a Video Image Processor (VIP) are presented. Approaches to the enhancement of the performance of these algorithms with knowledge-based techniques and the potential for deployment of highly-parallel multi-processor systems for these algorithms are discussed.

Symosek, P.↗

Numerical Modeling of Turbulence Effects within an Evaporating Droplet in Atomizing Sprays

A new approach to account for finite thermal conductivity and turbulence effects within atomizing liquid sprays is presented in this paper. The model is an extension of the T-blob and T-TAB atomization/spray model of Trinh and Chen (2005). This finite conductivity model is based on the two-temperature film theory, where the turbulence characteristics of the droplet are used to estimate the effective thermal diffhsivity within the droplet phase. Both one-way and two-way coupled calculations were performed to investigate the performance of this model. The current evaporation model is incorporated into the T-blob atomization model of Trinh and Chen (2005) and implemented in an existing CFD Eulerian-Lagrangian two-way coupling numerical scheme. Validation studies were carried out by comparing with available evaporating atomization spray experimental data in terms of jet penetration, temperature field, and droplet SMD distribution within the spray. Validation results indicate the superiority of the finite-conductivity model in low speed parallel flow evaporating spray.

Balasubramanyam, M. S.↗

Blade row dynamic digital compressor program. Volume 1: J85 clean inlet flow and parallel compressor models

The results are presented of a one-dimensional dynamic digital blade row compressor model study of a J85-13 engine operating with uniform and with circumferentially distorted inlet flow. Details of the geometry and the derived blade row characteristics used to simulate the clean inlet performance are given. A stability criterion based upon the self developing unsteady internal flows near surge provided an accurate determination of the clean inlet surge line. The basic model was modified to include an arbitrary extent multi-sector parallel compressor configuration for investigating 180 deg 1/rev total pressure, total temperature, and combined total pressure and total temperature distortions. The combined distortions included opposed, coincident, and 90 deg overlapped patterns. The predicted losses in surge pressure ratio matched the measured data trends at all speeds and gave accurate predictions at high corrected speeds where the slope of the speed lines approached the vertical.

Tesch, W. A.↗

Programmable remapper with single flow architecture

The invention relates to image processing systems and methods and in particular to a machine which accepts a real time video image in the form of a matrix of picture elements (pixels) and remaps such image according to a selectable one of a plurality of mapping functions to create an output matrix of pixels. Such mapping functions, or transformations, may be any one of a number of different transformations depending on the objective of the user of the system. The system remaps input images from one coordinate system to another using a set of look-up tables for the data necessary for the transform. The transforms, which are operator selectable, are precomputed and loaded into massive look-up tables. Input pixels, via the look-up tables of any particular transform selected, are mapped into output pixels with the radiance information of the input pixels being appropriately weighted. An earlier embodiment of the system included two parallel processors: a collective processor which mapped multiple input pixels into a single output pixel and an interpolative processor. The interpolative processor performed an interpolation among pixels in the input image where a given input pixel may affect the value of many output pixels. Several advantages are provided over previous embodiments in that the two distinct processors are replaced by a single processor capable of performing both types of operations (collective and interpolative) with no more complexity. Previously, there has existed no image processor or 'remapper' that can operate with sufficient speed and flexibility to permit investigating different transformation patterns in real time.

Fisher, Timothy E.↗