Search NASASearch

SEARCH · Search NASA

Results for “Computer systems performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Experimental Evaluation and Workload Characterization for High-Performance Computer Architectures

This research is conducted in the context of the Joint NSF/NASA Initiative on Evaluation (JNNIE). JNNIE is an inter-agency research program that goes beyond typical.bencbking to provide and in-depth evaluations and understanding of the factors that limit the scalability of high-performance computing systems. Many NSF and NASA centers have participated in the effort. Our research effort was an integral part of implementing JNNIE in the NASA ESS grand challenge applications context. Our research work under this program was composed of three distinct, but related activities. They include the evaluation of NASA ESS high- performance computing testbeds using the wavelet decomposition application; evaluation of NASA ESS testbeds using astrophysical simulation applications; and developing an experimental model for workload characterization for understanding workload requirements. In this report, we provide a summary of findings that covers all three parts, a list of the publications that resulted from this effort, and three appendices with the details of each of the studies using a key publication developed under the respective work.

El-Ghazawi, Tarek A.

Characterization Study of TestBed Infrastructure Performance in a Distributed Simulation Environment: Baseline Analysis

Characterization of the performance of Air Traffic Management Exploration (ATM-X) TestBed integration environment has been investigated and documented for one system configuration for progressively increasing traffic. Several statistical parameters were used to assess the performance of the TestBed distributed system such as mean, standard deviation, skewness, and kurtosis of latency, and update rate for aircraft state messages that are transmitted through the simulated system under investigation. It is necessary to assess the performance characteristics of distributed systems in terms of the indicated statistical parameters mentioned above. It is critical to verify the system performance with respect to a researcher’s required system performance. Computer host specifications are documented in terms of Central Processing Unit (CPU) clock speed and core count. Transmission Control Protocol/ Internet Protocol (TCP/IP) message protocol was used for data transmission. The system network topology also contributes to the latency and update rate variations from the one imposed by the data source. The motivation for selecting the TestBed infrastructure as the focus of this study can be attributed to the number of services and capabilities it provides that help simplify the process of preparing and conducting a simulation. These capabilities include an easy to use GUI for simulation configuration, access to TestBed library by the end-user of other simulation software components, a modular adapter paradigm that allows simple connectivity of external software to TestBed, connectivity with other simulation laboratories, and a Software Development Kit (SDK) for quicker development. Two types of traffic generators, Air Traffic Generator (ATG) and Multi Aircraft Control System (MACS) were used to generate messages that were injected into the TestBed distributed environment. Eight different air traffic scenarios with progressively increasing loads were generated for each air traffic simulator. The corresponding air traffic loads between the two simulators had an identical number of aircraft per scenario, but different flight plans. It was observed that the performance of MACS degraded for air traffic scenarios containing more than 200 aircraft (37.5 KB/s nominal throughput). However, ATG performed adequately under all tested air traffic loads up to 1200 aircraft (225. KB/s nominal throughput). The tests show that MACS exhibits better latency performance with smaller aircraft loads when compared to ATG. The tests also show that the TestBed infrastructure successfully transmits 1200 aircraft without significant degradation of its performance. From the latency trends for both MACS and ATG, it is clear that as aircraft load increases, the latency in the system increases as well as its standard deviation. Likewise, the trends for the update data rate for both MACS and ATG show that as the aircraft load increases, so does the standard deviation and mean of the update rates which can be attributed to the performance of MACS and ATG applications. The analysis of the results of this study have proven that the overall system performance is dependent on the individual performance of each system component that is connected to TestBed, which subsequently propagates into the system. All TestBed characterization tests were conducted in SimLabs at NASA Ames Research Center in November 2019. This study addresses the need for a baseline TestBed characterization, and the results will serve as a reference for more complex simulation systems.

Air Traffic Management simulations

Overview of the NASA Glenn Flux Reconstruction Based High-Order Unstructured Grid Code

A computational fluid dynamics code based on the flux reconstruction (FR) method is currently being developed at NASA Glenn Research Center to ultimately provide a large- eddy simulation capability that is both accurate and efficient for complex aeropropulsion flows. The FR approach offers a simple and efficient method that is easy to implement and accurate to an arbitrary order on common grid cell geometries. The governing compressible Navier-Stokes equations are discretized in time using various explicit Runge-Kutta schemes, with the default being the 3-stage/3rd-order strong stability preserving scheme. The code is written in modern Fortran (i.e., Fortran 2008) and parallelization is attained through MPI for execution on distributed-memory high-performance computing systems. An h- refinement study of the isentropic Euler vortex problem is able to empirically demonstrate the capability of the FR method to achieve super-accuracy for inviscid flows. Additionally, the code is applied to the Taylor-Green vortex problem, performing numerous implicit large-eddy simulations across a range of grid resolutions and solution orders. The solution found by a pseudo-spectral code is commonly used as a reference solution to this problem, and the FR code is able to reproduce this solution using approximately the same grid resolution. Finally, an examination of the code's performance demonstrates good parallel scaling, as well as an implementation of the FR method with a computational cost/degree- of-freedom/time-step that is essentially independent of the solution order of accuracy for structured geometries.

High-Order Methods

Preliminary results from the NASA general aviation demonstration advanced avionics system program

NASA's Demonstration Advanced Avionics System (DAAS) is an integrated avionics system employing microprocessor technologies, data busing, and shared electronics displays. A DAAS demonstration system has been assessed through flight testing and demonstration for potential users which includes among its functions autopiloting, navigation/flight planning, a flight warning and advisory system, performance computations, normal and emergency checklists, and a ground simulation function. Exceptional performance has been obtained from the DAAS electronic horizontal situation indicator, the autopilot, the navigator/flight planner, and the discrete address beacon system.

Hardy, G. H.

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry

Computational Performance of Progressive Damage Analysis of Composite Laminates using Abaqus/Explicit with 16 to 512 CPU Cores

The computational scaling performance of progressive damage analysis using Abaqus/ Explicit is evaluated and quantified using from 16 to 512 CPU cores. Several analyses were conducted on varying numbers of cores to determine the scalability of the code on five NASA high performance computing systems. Two finite element models representative of typical models used for progressive damage analysis of composite laminates were used. The results indicate a 10 to 15 times speed up scaling from 24 to 512 cores. The run times were modestly reduced with newer generations of CPU hardware. If the number of degrees of freedom is held constant with respect to the number of cores, the model size can be increased by a factor of 20, scaling from 16 to 512 cores, with the same run time. An empirical expression was derived relating run time, the number of cores, and the number of degrees of freedom. Analysis cost was examined in terms of software tokens and hardware utilization. Using additional cores reduces token usage since the computational performance increases more rapidly than the token requirement with increasing number of cores. The in- crease in hardware cost with increasing cores was found to be modest. Overall the results show relatively good scalability of the Abaqus/Explicit code on up to 512 cores.

Bergan, A. C.

A Unified Air-Sea Visualization System: Survey on Gridding Structures

The goal is to develop a Unified Air-Sea Visualization System (UASVS) to enable the rapid fusion of observational, archival, and model data for verification and analysis. To design and develop UASVS, modelers were polled to determine the gridding structures and visualization systems used, and their needs with respect to visual analysis. A basic UASVS requirement is to allow a modeler to explore multiple data sets within a single environment, or to interpolate multiple datasets onto one unified grid. From this survey, the UASVS should be able to visualize 3D scalar/vector fields; render isosurfaces; visualize arbitrary slices of the 3D data; visualize data defined on spectral element grids with the minimum number of interpolation stages; render contours; produce 3D vector plots and streamlines; provide unified visualization of satellite images, observations and model output overlays; display the visualization on a projection of the users choice; implement functions so the user can derive diagnostic values; animate the data to see the time-evolution; animate ocean and atmosphere at different rates; store the record of cursor movement, smooth the path, and animate a window around the moving path; repeatedly start and stop the visual time-stepping; generate VHS tape animations; work on a variety of workstations; and allow visualization across clusters of workstations and scalable high performance computer systems.

Anand, Harsh

Automatic Monitoring Of Complicated Systems

Collection of computer programs developed for expert computer system performing complicated, tedious, and repetitive portions of analysis of telemetry data from spacecraft. Provides nonstop, accurate surveillance of incoming data, also frees operators to concentrate their expertise on unexpected abnormal operating conditions. When unable to explain discrepancies with certainty resulting from data out of synchronization or other falso-alarm conditions, triggers alarm devices to request assistance from designated individuals. Concept useful in such terrestrial systems as production lines, power-distribution networks, chemical processes, large airplanes, and other assemblies of interdependent equipment.

Schwuttke, Ursula M.

Requirements for multidisciplinary design of aerospace vehicles on high performance computers

The design of aerospace vehicles is becoming increasingly complex as the various contributing disciplines and physical components become more tightly coupled. This coupling leads to computational problems that will be tractable only if significant advances in high performance computing systems are made. Some of the modeling, algorithmic and software requirements generated by the design problem are discussed.

Voigt, Robert G.

High-performance equation solvers and their impact on finite element analysis

The role of equation solvers in modern structural analysis software is described. Direct and iterative equation solvers which exploit vectorization on modern high-performance computer systems are described and compared. The direct solvers are two Cholesky factorization methods. The first method utilizes a novel variable-band data storage format to achieve very high computation rates and the second method uses a sparse data storage format designed to reduce the number of operations. The iterative solvers are preconditioned conjugate gradient methods. Two different preconditioners are included; the first uses a diagonal matrix storage scheme to achieve high computation rates and the second requires a sparse data storage scheme and converges to the solution in fewer iterations that the first. The impact of using all of the equation solvers in a common structural analysis software system is demonstrated by solving several representative structural analysis problems.

Poole, Eugene L.

High-performance equation solvers and their impact on finite element analysis

The role of equation solvers in modern structural analysis software is described. Direct and iterative equation solvers which exploit vectorization on modern high-performance computer systems are described and compared. The direct solvers are two Cholesky factorization methods. The first method utilizes a novel variable-band data storage format to achieve very high computation rates and the second method uses a sparse data storage format designed to reduce the number od operations. The iterative solvers are preconditioned conjugate gradient methods. Two different preconditioners are included; the first uses a diagonal matrix storage scheme to achieve high computation rates and the second requires a sparse data storage scheme and converges to the solution in fewer iterations that the first. The impact of using all of the equation solvers in a common structural analysis software system is demonstrated by solving several representative structural analysis problems.

Poole, Eugene L.

Linear analysis of opto-mechanical systems

A general framework is presented for the matrix-form, linear optical model analysis of controlled optomechanical systems; the models are conjoined with linear models of structures and controls to compute system performance as a function of optics, structures, and control parameters. Covariance analysis, optimization, and estimation/simulation are used. Attention is given to a tolerancing example for the Hubble Space Telescope's Wide Field and Planetary Camera, which involves the creation of a linear model of residual pupil shear.

Redding, David C.

Developments in Cylindrical Shell Stability Analysis

Today high-performance computing systems and new analytical and numerical techniques enable engineers to explore the use of advanced materials for shell design. This paper reviews some of the historical developments of shell buckling analysis and design. The paper concludes by identifying key research directions for reliable and robust methods development in shell stability analysis and design.

Knight, Norman F., Jr.

NAS Parallel Benchmarks I/O Version 2.4

We describe a benchmark problem, based on the Block-Tridiagonal (BT) problem of the NAS Parallel Benchmarks (NPB), which is used to test the output capabilities of high-performance computing systems, especially parallel systems. We also present a source code implementation of the benchmark, called NPBIO2.4-MPI, based on the MPI implementation of NPB, using a variety of ways to write the computed solutions to file.

Wong, Parkson

Analyzing Contents of a Computer Cache

The Cache Contents Estimator (CCE) is a computer program that provides information on the contents of level-1 cache of a PowerPC computer. The CCE is configurable to enable simulation of any processor in the PowerPC family. The need for CCE arises because the contents of level-1 caches are not available to either hardware or software readout mechanisms, yet information on the contents is crucial in the development of fault-tolerant or highly available computing systems and for realistic modeling and prediction of computing- system performance. The CCE comprises two independent subprograms: (1) the Dynamic Application Address eXtractor (DAAX), which extracts the stream of address references from an application program undergoing execution and (2) the Cache Simulator (CacheSim), which models the level-1 cache of the processor to be analyzed, by mimicking what the cache controller would do, in response to the address stream from DAAX. CacheSim generates a running estimate of the contents of the data and the instruction subcaches of the level-1 cache, hit/miss ratios, the percentage of cache that contains valid or active data, and time-stamped histograms of the cache content.

Beahan, John

Climatespark: an In-Memory Distributed Computing Framework for Big Climate Data Analytics

The unprecedented growth of climate data creates new opportunities for climate studies, and yet big climate data pose a grand challenge to climatologists to efficiently manage and analyze big data. The complexity of climate data content and analytical algorithms increases the difficulty of implementing algorithms on high performance computing systems. This paper proposes an in-memory, distributed computing framework, ClimateSpark, to facilitate complex big data analytics and time-consuming computational tasks. Chunking data structure improves parallel I/O efficiency, while a spatiotemporal index is built for the chunks to avoid unnecessary data reading and preprocessing. An integrated, multi-dimensional, array-based data model (ClimateRDD) and ETL operations are developed to address big climate data variety by integrating the processing components of the climate data lifecycle. ClimateSpark utilizes Spark SQL and Apache Zeppelin to develop a web portal to facilitate the interaction among climatologists, climate data, analytic operations and computing resources (e.g., using SQL query and Scala/Python notebook). Experimental results show that ClimateSpark conducts different spatiotemporal data queries/analytics with high efficiency and data locality. ClimateSpark is easily adaptable to other big multiple- dimensional, array-based datasets in various geoscience domains.

Hu, Fei

Optical Error Budgeting Using Linearized Ray-Trace Models

The Root-Sum-Squared, or “RSS” wavefront error model is a simple, scalar tool, commonly used for space telescope error budgeting. At the same time, much more detailed models, combining ray-trace and Fourier optics with optical alignments and wavefront controls, can provide accurate, high -resolution simulations for detailed system and subsystem design. This paper makes a connection between the two modeling approaches by deriving RSS model coefficients from ray-trace models, including the effects of wavefront controls, for computing system performance from component error statistics. It is shown that, properly constructed, the simple RSS error budget is a covariance analysis, and can be as accurate as high-resolution wavefront models for statistical wavefront error prediction. A notional segmented-aperture space telescope is used to illustrate this error modeling process.

Wavefront error