Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Visualization of Unsteady Computational Fluid Dynamics

The current compute environment that most researchers are using for the calculation of 3D unsteady Computational Fluid Dynamic (CFD) results is a super-computer class machine. The Massively Parallel Processors (MPP's) such as the 160 node IBM SP2 at NAS and clusters of workstations acting as a single MPP (like NAS's SGI Power-Challenge array and the J90 cluster) provide the required computation bandwidth for CFD calculations of transient problems. If we follow the traditional computational analysis steps for CFD (and we wish to construct an interactive visualizer) we need to be aware of the following: (1) Disk space requirements. A single snap-shot must contain at least the values (primitive variables) stored at the appropriate locations within the mesh. For most simple 3D Euler solvers that means 5 floating point words. Navier-Stokes solutions with turbulence models may contain 7 state-variables. (2) Disk speed vs. Computational speeds. The time required to read the complete solution of a saved time frame from disk is now longer than the compute time for a set number of iterations from an explicit solver. Depending, on the hardware and solver an iteration of an implicit code may also take less time than reading the solution from disk. If one examines the performance improvements in the last decade or two, it is easy to see that depending on disk performance (vs. CPU improvement) may not be the best method for enhancing interactivity. (3) Cluster and Parallel Machine I/O problems. Disk access time is much worse within current parallel machines and cluster of workstations that are acting in concert to solve a single problem. In this case we are not trying to read the volume of data, but are running the solver and the solver outputs the solution. These traditional network interfaces must be used for the file system. (4) Numerics of particle traces. Most visualization tools can work upon a single snap shot of the data but some visualization tools for transient problems require dealing with time.

Haimes, Robert↗

Parallel Visualization Co-Processing of Overnight CFD Propulsion Applications

An interactive visualization system pV3 is being developed for the investigation of advanced computational methodologies employing visualization and parallel processing for the extraction of information contained in large-scale transient engineering simulations. Visual techniques for extracting information from the data in terms of cutting planes, iso-surfaces, particle tracing and vector fields are included in this system. This paper discusses improvements to the pV3 system developed under NASA's Affordable High Performance Computing project.

Edwards, David E.↗

Memory-Intensive Benchmarks: IRAM vs. Cache-Based Machines

The increasing gap between processor and memory performance has lead to new architectural models for memory-intensive applications. In this paper, we explore the performance of a set of memory-intensive benchmarks and use them to compare the performance of conventional cache-based microprocessors to a mixed logic and DRAM processor called VIRAM. The benchmarks are based on problem statements, rather than specific implementations, and in each case we explore the fundamental hardware requirements of the problem, as well as alternative algorithms and data structures that can help expose fine-grained parallelism or simplify memory access patterns. The benchmarks are characterized by their memory access patterns, their basic control structures, and the ratio of computation to memory operation.

Biswas, Rupak↗

Revised Calibration Strategy for the CALIOP 532 nm Channel: Daytime - Part II

The CALIPSO lidar (CALIOP) makes backscatter measurements at 532 nm and 1064 nm and linear depolarization ratios at 532 nm. Accurate calibration of the backscatter measurements is essential in the retrieval of optical properties. An assessment of the nighttime 532 nm parallel channel calibration showed that the calibration strategy used for the initial release (Release 1) of the CALIOP lidar level 1B data was acceptable. In general, the nighttime calibration coefficients are relatively constant over the darkest segment of the orbit, but then change rapidly over a short period as the satellite enters sunlight. The daytime 532 nm parallel channel calibration scheme implemented in Release 1 derived the daytime calibration coefficients from the previous nighttime coefficients. A subsequent review of the daytime 532 nm parallel channel calibration revealed that the daytime calibration coefficients do not remain constant, but vary considerably over the course of the orbit, due to thermally-induced misalignment of the transmitter and receiver. A correction to the daytime calibration scheme is applied in Release 2 of the data. Results of both nighttime and daytime calibration performance are presented in this paper.

Powell, Kathleen A.↗

A Simple GPU-Accelerated Two-Dimensional MUSCL-Hancock Solver for Ideal Magnetohydrodynamics

We describe our experience using NVIDIA's CUDA (Compute Unified Device Architecture) C programming environment to implement a two-dimensional second-order MUSCL-Hancock ideal magnetohydrodynamics (MHD) solver on a GTX 480 Graphics Processing Unit (GPU). Taking a simple approach in which the MHD variables are stored exclusively in the global memory of the GTX 480 and accessed in a cache-friendly manner (without further optimizing memory access by, for example, staging data in the GPU's faster shared memory), we achieved a maximum speed-up of approx. = 126 for a sq 1024 grid relative to the sequential C code running on a single Intel Nehalem (2.8 GHz) core. This speedup is consistent with simple estimates based on the known floating point performance, memory throughput and parallel processing capacity of the GTX 480.

graphics processing units↗

The TEKLIB graphic library

TEKLIB is a library of procedures written in TI PASCAL to perform basic graphic tasks. TEKLIB was written to provide an interface between a graphics terminal and the TI 990. The TI 990 is used as a controller for the Finite Element Machine which is an array of microprocessors designed to solve problems by finite element methods in parallel. The use of TEKLIB provides a means of inputting data graphically and displaying output.

Bostic, S. W.↗

Unstructured grids on SIMD torus machines

Unstructured grids lead to unstructured communication on distributed memory parallel computers, a problem that has been considered difficult. Here, we consider adaptive, offline communication routing for a SIMD processor grid. Our approach is empirical. We use large data sets drawn from supercomputing applications instead of an analytic model of communication load. The chief contribution of this paper is an experimental demonstration of the effectiveness of certain routing heuristics. Our routing algorithm is adaptive, nonminimal, and is generally designed to exploit locality. We have a parallel implementation of the router, and we report on its performance.

Bjorstad, Petter E.↗

Latency Hiding in Dynamic Partitioning and Load Balancing of Grid Computing Applications

The Information Power Grid (IPG) concept developed by NASA is aimed to provide a metacomputing platform for large-scale distributed computations, by hiding the intricacies of highly heterogeneous environment and yet maintaining adequate security. In this paper, we propose a latency-tolerant partitioning scheme that dynamically balances processor workloads on the.IPG, and minimizes data movement and runtime communication. By simulating an unsteady adaptive mesh application on a wide area network, we study the performance of our load balancer under the Globus environment. The number of IPG nodes, the number of processors per node, and the interconnected speeds are parameterized to derive conditions under which the IPG would be suitable for parallel distributed processing of such applications. Experimental results demonstrate that effective solution are achieved when the IPG nodes are connected by a high-speed asynchronous interconnection network.

Das, Sajal K.↗

Traffic Flow Management and Optimization

This talk will present an overview of Traffic Flow Management (TFM) research at NASA Ames Research Center. Dr. Rios will focus on his work developing a large-scale, parallel approach to solving traffic flow management problems in the national airspace. In support of this talk, Dr. Rios will provide some background on operational aspects of TFM as well a discussion of some of the tools needed to perform such work including a high-fidelity airspace simulator. Current, on-going research related to TFM data services in the national airspace system and general aviation will also be presented.

linear programming↗

Processing CCD Images to Detect Transits of Earth-Sized Planets: Maximizing Sensitivity While Achieving Reasonable Downlink Requirements

We have performed end-to-end laboratory and numerical simulations to demonstrate the capability of differential photometry under realistic operating conditions to detect transits of Earth-sized planets orbiting solar-like stars. Data acquisition and processing were conducted using the same methods planned for the proposed Kepler Mission. These included performing aperture photometry on large-format CCD images of an artificial star fields obtained without a shutter at a readout rate of 1 megapixel/sec, detecting and removing cosmic rays from individual exposures and making the necessary corrections for nonlinearity and shutterless operation in the absence of darks. We will discuss the image processing tasks performed `on-board' the simulated spacecraft, which yielded raw photometry and ancillary data used to monitor and correct for systematic effects, and the data processing and analysis tasks conducted to obtain lightcurves from the raw data and characterize the detectability of transits. The laboratory results are discussed along with the results of a numerical simulation carried out in parallel with the laboratory simulation. These two simulations demonstrate that a system-level differential photometric precision of 10-5 on five- hour intervals can be achieved under realistic conditions.

Earth-size planets↗

Turbine Electrified Energy Management for Single Aisle Aircraft

Electrified aircraft propulsion technology is being developed to reduce the environmental impacts of the aviation industry. This is prompting the exploration of potential uses and benefits of hybrid systems in which electric powertrains are integrated with more traditional gas turbine propulsion systems. Turbine Electrified Energy Management (TEEM) is an energy management approach for hybrid-electric architectures in which electric machines are connected to the turbofan shafts and used to suppress the off-design operation naturally associated with engine transients. This reduces the need to maintain a large amount of compressor operability margin, thus allowing further exploration of the engine design space. In this study, a 19,000 lbf engine within a parallel hybrid propulsion system is considered along with a 30,000 lbf standalone engine. Data from prior TEEM applications are used to approximate the electric machine sizing required to achieve operability benefits. The TEEM controller is shown to improve operability during transients through the reduction of stall margin undershoots and the decrease of transient variations in component performance maps by over 29%.

Controls↗

Turbine Electrified Energy Management for Single Aisle Aircraft

Electrified aircraft propulsion technology is being developed to reduce the environmental impacts of the aviation industry. This is prompting the exploration of potential uses and benefits of hybrid systems in which electric powertrains are integrated with more traditional gas turbine propulsion systems. Turbine Electrified Energy Management (TEEM) is an energy management approach for hybrid-electric architectures in which electric machines are connected to the turbofan shafts and used to suppress the off-design operation naturally associated with engine transients. This reduces the need to maintain a large amount of compressor operability margin, thus allowing further exploration of the engine design space. In this study, a 19,000 lbf engine within a parallel hybrid propulsion system is considered along with a 30,000 lbf standalone engine. Data from prior TEEM applications are used to approximate the electric machine sizing required to achieve operability benefits. The TEEM controller is shown to improve operability during transients through the reduction of stall margin undershoots and the decrease of transient variations in component performance maps by over 29%.

EAP↗

Liquid Nitrogen (Oxygen Simulent) Thermodynamic Venting System Test Data Analysis

In designing systems for the long-term storage of cryogens in low gravity space environments, one must consider the effects of thermal stratification on excessive tank pressure that will occur due to environmental heat leakage. During low gravity operations, a Thermodynamic Venting System (TVS) concept is expected to maintain tank pressure without propellant resettling. The TVS consists of a recirculation pump, Joule-Thomson (J-T) expansion valve, and a parallel flow concentric tube heat exchanger combined with a longitudinal spray bar. Using a small amount of liquid extracted by the pump and passing it though the J-T valve, then through the heat exchanger, the bulk liquid and ullage are cooled, resulting in lower tank pressure. A series of TVS tests were conducted at the Marshall Space Flight Center using liquid nitrogen as a liquid oxygen simulant. The tests were performed at fill levels of 90%, 50%, and 25% with gaseous nitrogen and helium pressurants, and with a tank pressure control band of 7 kPa. A transient one-dimensional model of the TVS is used to analyze the data. The code is comprised of four models for the heat exchanger, the spray manifold and injector tubes, the recirculation pump, and the tank. The TVS model predicted ullage pressure and temperature and bulk liquid saturation pressure and temperature are compared with data. Details of predictions and comparisons with test data regarding pressure rise and collapse rates will be presented in the final paper.

Hedayat, A.↗

Software Engineering Support of the Third Round of Scientific Grand Challenge Investigations: Earth System Modeling Software Framework Survey

One of the most significant challenges in large-scale climate modeling, as well as in high-performance computing in other scientific fields, is that of effectively integrating many software models from multiple contributors. A software framework facilitates the integration task, both in the development and runtime stages of the simulation. Effective software frameworks reduce the programming burden for the investigators, freeing them to focus more on the science and less on the parallel communication implementation. while maintaining high performance across numerous supercomputer and workstation architectures. This document surveys numerous software frameworks for potential use in Earth science modeling. Several frameworks are evaluated in depth, including Parallel Object-Oriented Methods and Applications (POOMA), Cactus (from (he relativistic physics community), Overture, Goddard Earth Modeling System (GEMS), the National Center for Atmospheric Research Flux Coupler, and UCLA/UCB Distributed Data Broker (DDB). Frameworks evaluated in less detail include ROOT, Parallel Application Workspace (PAWS), and Advanced Large-Scale Integrated Computational Environment (ALICE). A host of other frameworks and related tools are referenced in this context. The frameworks are evaluated individually and also compared with each other.

Talbot, Bryan↗

What Multilevel Parallel Programs do when you are not Watching: A Performance Analysis Case Study Comparing MPI/OpenMP, MLP, and Nested OpenMP

With the current trend in parallel computer architectures towards clusters of shared memory symmetric multi-processors, parallel programming techniques have evolved that support parallelism beyond a single level. When comparing the performance of applications based on different programming paradigms, it is important to differentiate between the influence of the programming model itself and other factors, such as implementation specific behavior of the operating system (OS) or architectural issues. Rewriting-a large scientific application in order to employ a new programming paradigms is usually a time consuming and error prone task. Before embarking on such an endeavor it is important to determine that there is really a gain that would not be possible with the current implementation. A detailed performance analysis is crucial to clarify these issues. The multilevel programming paradigms considered in this study are hybrid MPI/OpenMP, MLP, and nested OpenMP. The hybrid MPI/OpenMP approach is based on using MPI [7] for the coarse grained parallelization and OpenMP [9] for fine grained loop level parallelism. The MPI programming paradigm assumes a private address space for each process. Data is transferred by explicitly exchanging messages via calls to the MPI library. This model was originally designed for distributed memory architectures but is also suitable for shared memory systems. The second paradigm under consideration is MLP which was developed by Taft. The approach is similar to MPi/OpenMP, using a mix of coarse grain process level parallelization and loop level OpenMP parallelization. As it is the case with MPI, a private address space is assumed for each process. The MLP approach was developed for ccNUMA architectures and explicitly takes advantage of the availability of shared memory. A shared memory arena which is accessible by all processes is required. Communication is done by reading from and writing to the shared memory.

Jost, Gabriele↗

Assessing the Ability of Instantaneous Aircraft and Sonde Measurements to Characterize Climatological Means and Long-Term Trends in Tropospheric Composition

Over four decades of measurements exist that sample the 3-D composition of reactive trace gases in the troposphere from approximately weekly ozone sondes, instrumentation on civil aircraft, and individual comprehensive aircraft field campaigns. An obstacle to using these data to evaluate coupled chemistry-climate models (CCMs)the models used to project future changes in atmospheric composition and climateis that exact space-time matching between model fields and observations cannot be done, as CCMs generate their own meteorology. Evaluation typically involves averaging over large spatiotemporal regions, which may not reflect a true average due to limited or biased sampling. This averaging approach generally loses information regarding specific processes. Here we aim to identify where discrete sampling may be indicative of long-term mean conditions, using the GEOS-Chem global chemical-transport model (CTM) driven by the MERRA reanalysis to reflect historical meteorology from 2003 to 2012 at 2o by 2.5o resolution. The model has been sampled at the time and location of every ozone sonde profile available from the Would Ozone and Ultraviolet Radiation Data Centre (WOUDC), along the flight tracks of the IAGOSMOZAICCARABIC civil aircraft campaigns, as well as those from over 20 individual field campaigns performed by NASA, NOAA, DOE, NSF, NERC (UK), and DLR (Germany) during the simulation period. Focusing on ozone, carbon monoxide and reactive nitrogen species, we assess where aggregates of the in situ data are representative of the decadal mean vertical, spatial and temporal distributions that would be appropriate for evaluating CCMs. Next, we identically sample a series of parallel sensitivity simulations in which individual emission sources (e.g., lightning, biogenic VOCs, wildfires, US anthropogenic) have been removed one by one, to assess where and when the aggregated observations may offer constraints on these processes within CCMs. Lastly, we show results of an additional 31-year simulation from 1980-2010 of GEOS-Chem driven by the MACCity emissions inventory and MERRA reanalysis at 4o by 5o. We sample the model at every WOUDC sonde and flight track from MOZAIC and NASA field campaigns to evaluate which aggregate observations are statistically reflective of long-term trends over the period.

Atmospheric composition↗

Analysis of the Value Added When Deploying a Model-Based Approach for the Validation and Verification of the Medical Database Software

The Medical Database (MD) is a virtual repository consisting of two software components: Medical Item Database (MedID) and the Evidence Library (EL). MedID consists of engineering data and associated information for specific medical resource items (e.g., pharmaceutical, medical devices, and supporting components), while the EL is a tool which provides all of the medical evidence necessary. The MD will 1) serve as the single “source of truth” for the Informing Mission Planning via Analysis of Complex Tradespaces (IMPACT) tool suite for both medical evidence and medical resource engineering data and 2) will be used in conjunction with the IMPACT tool suite to inform research prioritizations and perform systematic trade study evaluations to aid stakeholders in making informed decisions regarding simulated human spaceflight missions. The MD project used a Model-Based Systems Engineering (MBSE) approach to support all life cycles of the software development, while in parallel the human factors engineering team used modeling to support Human Centered Design (HCD) strategies in an effort to improve software usability. HCD is a frequently used approach in design frameworks that develops resolutions to complexities and challenges by involving the human perspective in all steps of the problem-solving process. By integrating the model-based approaches used for systems engineering and human factors activities, the project is able to leverage the model-based artifacts originally created for HCD activities for system level and human factors validation. In this presentation, our team highlights the value added when leveraging these model-based artifacts to support the on-going verification and validation activities.

C. Laing↗

A real time neural net estimator of fatigue life

A neural net architecture is proposed to estimate, in real-time, the fatigue life of mechanical components, as part of the Intelligent Control System for Reusable Rocket Engines. Arbitrary component loading values were used as input to train a two hidden-layer feedforward neural net to estimate component fatigue damage. The ability of the net to learn, based on a local strain approach, the mapping between load sequence and fatigue damage has been demonstrated for a uniaxial specimen. Because of its demonstrated performance, the neural computation may be extended to complex cases where the loads are biaxial or triaxial, and the geometry of the component is complex (e.g., turbopump blades). The generality of the approach is such that load/damage mappings can be directly extracted from experimental data without requiring any knowledge of the stress/strain profile of the component. In addition, the parallel network architecture allows real-time life calculations even for high frequency vibrations. Owing to its distributed nature, the neural implementation will be robust and reliable, enabling its use in hostile environments such as rocket engines. This neural net estimator of fatigue life is seen as the enabling technology to achieve component life prognosis, and therefore would be an important part of life extending control for reusable rocket engines.

Troudet, T.↗