Search NASASearch

SEARCH · Search NASA

Results for “processor performance factors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

X2000 advanced avionics characterization study

The characterization study has shown that adjustments in an application's data accesses can easily create a performance difference of three or more times in actual applications. In general, by iterating over a small portion of a data set rather than its entirety, execution can remain within cache thereby producing a performance increase. Additionally, the approaches used in performing I/O can make a major performance difference.

processor performance factors

A modified reconfigurable data path processor

High throughput is an overriding factor dictating system performance. A configurable data processor is presented which can be modified to optimize performance for a wide class of problems. The new processor is specifically designed for arbitrary data path operations and can be dynamically reconfigured.

Ganesh, G.

Performance limitations in parallel processor simulations

A jet-engine model is partitioned and simulated on a parallel processor system consisting of five 8086/8087 floating-point computers. The simulation uses Heun's integration method. A near-optimal parallel simulation (in the sense of minimum execution time) achieves speedup of only 2.13 and efficiency of 42.6 percent, in effect wasting 57.4 percent of the available processing power. A detailed analysis identifies and graphically demonstrates why the system fails to achieve ideal performance (viz., speedup of 5 and efficiency of 100 percent). Inherent characteristics of the problem equations and solution algorithm account for the loss of nearly half of the available processing power. Overheads associated with interprocessor communication and processor synchronization account for only a small fraction of the lost processing power. The effects of these and other factors which limit parallel processor performance are illustrated through real-time timing-analyzer tracers describing the run/idle status of the parallel processors during the simulation.

O'Grady, E. Pearse

A simplified Integer Cosine Transform and its application in image compression

A simplified version of the integer cosine transform (ICT) is described. For practical reasons, the transform is considered jointly with the quantization of its coefficients. It differs from conventional ICT algorithms in that the combined factors for normalization and quantization are approximated by powers of two. In conventional algorithms, the normalization/quantization stage typically requires as many integer divisions as the number of transform coefficients. By restricting the factors to powers of two, these divisions can be performed by variable shifts in the binary representation of the coefficients, with speed and cost advantages to the hardware implementation of the algorithm. The error introduced by the factor approximations is compensated for in the inverse ICT operation, executed with floating point precision. The simplified ICT algorithm has potential applications in image-compression systems with disparate cost and speed requirements in the encoder and decoder ends. For example, in deep space image telemetry, the image processors on board the spacecraft could take advantage of the simplified, faster encoding operation, which would be adjusted on the ground, with high-precision arithmetic. A dual application is found in compressed video broadcasting. Here, a fast, high-performance processor at the transmitter would precompensate for the factor approximations in the inverse ICT operation, to be performed in real time, at a large number of low-cost receivers.

Costa, M.

Highly parallel sparse Cholesky factorization

Several fine grained parallel algorithms were developed and compared to compute the Cholesky factorization of a sparse matrix. The experimental implementations are on the Connection Machine, a distributed memory SIMD machine whose programming model conceptually supplies one processor per data element. In contrast to special purpose algorithms in which the matrix structure conforms to the connection structure of the machine, the focus is on matrices with arbitrary sparsity structure. The most promising algorithm is one whose inner loop performs several dense factorizations simultaneously on a 2-D grid of processors. Virtually any massively parallel dense factorization algorithm can be used as the key subroutine. The sparse code attains execution rates comparable to those of the dense subroutine. Although at present architectural limitations prevent the dense factorization from realizing its potential efficiency, it is concluded that a regular data parallel architecture can be used efficiently to solve arbitrarily structured sparse problems. A performance model is also presented and it is used to analyze the algorithms.

Gilbert, John R.

Design and Performance of the Astro-E/XRS Signal Processing System

We describe the signal processing system of the Astro-E XRS Instrument. The Calorimeter Analog Processor (CAP) provides bias and power for the detectors and amplifies the detector signals by a factor of 20,000. The Calorimeter Digital Processor (CDP) performs the digital processing of the calorimeter signals, detecting X-ray pulses and analyzing them by optimal filtering. We describe the operation of pulse detection, pulse height analysis, and risetime determination. We also discuss performance, including the three event grades (hi-res, mid-res, and low-res), anticoincidence detection, counting rate dependence, and noise rejection.

Boyce, K. R.

The Design and Performance of the Astro-E/XRS Signal Processing System

We describe the signal processing system of the Astro-E XRS instrument. The Calorimeter Analog Processor (CAP) provides bias and power for the detectors and amplifies the detector signals by a factor of 20,000. The Calorimeter Digital Processor (CDP) performs the digital processing of the calorimeter signals, detecting X-ray pulses and analyzing them by optimal filtering. We describe the operation of pulse detection, pulse height analysis, and risetime determination. We also discuss performance, including the three event grades (hi-res, mid-res, and low-res), anticoincidence detection, counting rate dependence, and noise rejection.

Boyce, K. R.

Design and Performance of the Astro-E/XRS Signal Processing System

We describe the signal processing system of the Astro-E XRS instrument. The Calorimeter Analog Processor (CAP) provides bias and power for the detectors and amplifies the detector signals by a factor of 20,000. The Calorimeter Digital Processor (CDP) performs the digital processing of the calorimeter signals, detecting X-ray pulses and analyzing them by optimal filtering. We describe the operation of pulse detection, Pulse height analysis. and risetime determination. We also discuss performance, including the three event grades (hi-res mid-res, and low-res). anticoincidence detection, counting rate dependence, and noise rejection.

Boyce, Kevin R.

An investigation of radiometer design using digital processing techniques

The use of digital signal processing techniques in Dicke switching radiometer design was investigated. The general approach was to develop an analytical model of the existing analog radiometer and identify factors which adversly affect its performance. A digital processor was then proposed to verify the feasibility of using digital techniques to minimize these adverse effects and improve the radiometer performance. Analysis and preliminary test results comparing the digital and analog processing approaches in radiometers design were analyzed.

Lawrence, R. W.

A Parallel Pipelined Renderer for the Time-Varying Volume Data

This paper presents a strategy for efficiently rendering time-varying volume data sets on a distributed-memory parallel computer. Time-varying volume data take large storage space and visualizing them requires reading large files continuously or periodically throughout the course of the visualization process. Instead of using all the processors to collectively render one volume at a time, a pipelined rendering process is formed by partitioning processors into groups to render multiple volumes concurrently. In this way, the overall rendering time may be greatly reduced because the pipelined rendering tasks are overlapped with the I/O required to load each volume into a group of processors; moreover, parallelization overhead may be reduced as a result of partitioning the processors. We modify an existing parallel volume renderer to exploit various levels of rendering parallelism and to study how the partitioning of processors may lead to optimal rendering performance. Two factors which are important to the overall execution time are re-source utilization efficiency and pipeline startup latency. The optimal partitioning configuration is the one that balances these two factors. Tests on Intel Paragon computers show that in general optimal partitionings do exist for a given rendering task and result in 40-50% saving in overall rendering time.

Chiueh, Tzi-Cker

Evaluation of Performance of Five Parallel Biological Water Proce

The objective of the work entitled Molecular Characterization of Eubacteria in a Biological Water Processor was to gain an understanding of the microbial diversity and species stability of the consortia that inhabit an anoxic bioreactor and to correlate those factors with functional performance, mechanical reliability, and stability. The evaluation was divided into four studies. During Study 1, replicate biological water processor (BWP) systems were operated to evaluate variability in the microbial diversity over time as a function of the initial consortia used for inoculation of the BWP reactors. Study 2 was designed to investigate the impact of an inoculum source on BWP performance. Study 3 was a modification of Study 2 where the impact of inoculum on BWP performance from inoculation until steady state operations was monitored. In Study 4, the reactors were divided into three different operational periods, based on the operational periods of the integrated water recovery test at the Johnson Space Center (JSC) in 2001.

Vega, Leticia M.

A performance study of sparse Cholesky factorization on INTEL iPSC/860

The problem of Cholesky factorization of a sparse matrix has been very well investigated on sequential machines. A number of efficient codes exist for factorizing large unstructured sparse matrices. However, there is a lack of such efficient codes on parallel machines in general, and distributed machines in particular. Some of the issues that are critical to the implementation of sparse Cholesky factorization on a distributed memory parallel machine are ordering, partitioning and mapping, load balancing, and ordering of various tasks within a processor. Here, we focus on the effect of various partitioning schemes on the performance of sparse Cholesky factorization on the Intel iPSC/860. Also, a new partitioning heuristic for structured as well as unstructured sparse matrices is proposed, and its performance is compared with other schemes.

Zubair, M.

Improved Dynamic Modeling of the Cascade Distillation Subsystem and Analysis of Factors Affecting Its Performance

The Cascade Distillation Subsystem (CDS) is a rotary multistage distiller being developed to serve as the primary processor for wastewater recovery during long-duration space missions. The CDS could be integrated with a system similar to the International Space Station Water Processor Assembly to form a complete water recovery system for future missions. A preliminary chemical process simulation was previously developed using Aspen Custom Modeler® (ACM), but it could not simulate thermal startup and lacked detailed analysis of several key internal processes, including heat transfer between stages. This paper describes modifications to the ACM simulation of the CDS that improve its capabilities and the accuracy of its predictions. Notably, the modified version can be used to model thermal startup and predicts the total energy consumption of the CDS. The simulation has been validated for both NaC1 solution and pretreated urine feeds and no longer requires retuning when operating parameters change. The simulation was also used to predict how internal processes and operating conditions of the CDS affect its performance. In particular, it is shown that the coefficient of performance of the thermoelectric heat pump used to provide heating and cooling for the CDS is the largest factor in determining CDS efficiency. Intrastage heat transfer affects CDS performance indirectly through effects on the coefficient of performance.

Perry, Bruce A.

Radiometric Compensation and Calibration for Radarsat ScanSAR

Due to lack of a standard for modeling the radar echo signal in terms of signal unit and coordinates as well as lack of a standard in designing the gain factors in each stage of a processor, absolute radiometric calibration of a SAR system is usually performed by treating the sensor and processor as one inseparable unit. This often makes the calibration procedure complicated and requiring the involvement of both radar system engineers and processor engineers in the whole process. This paper introduces a standard for modeling the radar echo signal and a standard in designing the gain factor of a ScanSAR processor. In this paper, the radar equation is derived based on the amount of energy instead of the power received from a backscatterer. These efforts lead to simple and easy-to-understand equations for radiometric compensation and calibration.

Jin, Michael Y.

Partial Overhaul and Initial Parallel Optimization of KINETICS, a Coupled Dynamics and Chemistry Atmosphere Model

KINETICS is a coupled dynamics and chemistry atmosphere model that is data intensive and computationally demanding. The potential performance gain from using a supercomputer motivates the adaptation from a serial version to a parallelized one. Although the initial parallelization had been done, bottlenecks caused by an abundance of communication calls between processors led to an unfavorable drop in performance. Before starting on the parallel optimization process, a partial overhaul was required because a large emphasis was placed on streamlining the code for user convenience and revising the program to accommodate the new supercomputers at Caltech and JPL. After the first round of optimizations, the partial runtime was reduced by a factor of 23; however, performance gains are dependent on the size of the data, the number of processors requested, and the computer used.

KINETICS

Internship Abstract and Final Reflection

The primary objective for this internship is the evaluation of an embedded natural language processor (NLP) as a way to introduce voice control into future space suits. An embedded natural language processor would provide an astronaut hands-free control for making adjustments to the environment of the space suit and checking status of consumables procedures and navigation. Additionally, the use of an embedded NLP could potentially reduce crew fatigue, increase the crewmember's situational awareness during extravehicular activity (EVA) and improve the ability to focus on mission critical details. The use of an embedded NLP may be valuable for other human spaceflight applications desiring hands-free control as well. An embedded NLP is unique because it is a small device that performs language tasks, including speech recognition, which normally require powerful processors. The dedicated device could perform speech recognition locally with a smaller form-factor and lower power consumption than traditional methods.

Sandor, Edward

Advanced technology for a satellite multichannel demultiplexer/demodulator

Satellite on-board processing is needed to efficiently service multiple users while at the same time minimizing earth station complexity. The processing satellite receives a wideband uplink at 30 GHz and down-converts it to a suitable intermediate frequency. A multichannel demultiplexer then separates the composite signal into discrete channels. Each channel is then demodulated by bulk demodulators, with the baseband signals routed to the downlink processor for retransmission to the receiving earth stations. This type of processing circumvents many of the difficulties associated with traditional bent-pipe repeater satellites. Uplink signal distortion and interference are not retransmitted on the downlink. Downlink power can be allocated in accordance with user needs, independent of uplink transmissions. This allows the uplink users to employ different data rates as well as different modulation and coding schemes. In addition, all downlink users have a common frequency standard and symbol clock on the satellite, which is useful for network synchronization in time division multiple access schemes. The purpose of this program is to demonstrate the concept of an optically implemented multichannel demultiplexer (MCD). A proof-of-concept (POC) model has been developed which has the ability to receive a 40 MHz wide composite signal consisting of up to 1000 40 kHz QPSK modulated channels and perform the demultiplexing process. In addition a set of special test equipment (STE) has been configured to evaluate the performance of the POC model. The optical MCD is realized as an acousto-optic spectrum analyzer utilizing the capability of Bragg cells to perform the required channelization. These Bragg cells receive an optical input from a laser source and an RF input (the signal). The Bragg interaction causes optical output diffractions at angles proportional to the RF input frequency. These discrete diffractions are optically detected and output to individual demodulators for baseband conversion. Optimization of the MCD design was conducted in order to achieve a compromise between two opposing sources of signal degradation: adjacent channel interference and intersymbol interference. The system was also optimized to allow simple, inexpensive ground stations communications with the MCD. These design goals led to the realization of a POC MCD which demonstrates the demultiplexing function with minimal signal degradation. Performance evaluation results using the STE equipment indicate that the dynamic range of the demultiplexer in the presence of adjacent and multiple channel loading is 40 - 50 dB. Measured bit error rate (BER) probabilities varied from the predicted theoretical results by one dB or less. The performance of the proof-of-concept model indicate that the development of a space qualified optically implemented MCD are feasible. The advantages to such an implementation include reduced size, weight and power and increased reliability when compared with electronic approaches. All of these factors are critical to on-board satellite processors. Further optimization can be conducted which trade ground station complexity and MCD performance to achieve desired system results.

Abramovitz, Irwin J.

Data management system performance modeling

This paper discusses analytical techniques that have been used to gain a better understanding of the Space Station Freedom's (SSF's) Data Management System (DMS). The DMS is a complex, distributed, real-time computer system that has been redesigned numerous times. The implications of these redesigns have not been fully analyzed. This paper discusses the advantages and disadvantages for static analytical techniques such as Rate Monotonic Analysis (RMA) and also provides a rationale for dynamic modeling. Factors such as system architecture, processor utilization, bus architecture, queuing, etc. are well suited for analysis with a dynamic model. The significance of performance measures for a real-time system are discussed.

Kiser, Larry M.