Search NASA⌕ Search

SEARCH · Search NASA

Results for “CPU”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A Multi-Architecture Approach for Implicit Computational Fluid Dynamics on Unstructured Grids

High-performance computing (HPC) architectures are trending toward manycore paradigms such as graphics processing units (GPUs). Approximately half of the top 100 publicly disclosed supercomputers in the world utilize GPU accelerators for performance. This is in contrast to a decade ago, where there were only a few such machines in the top 100. It is not currently possible to compile and run legacy central processing unit (CPU) software efficiently on GPUs without significant refactoring. Though a number of frameworks offering performance portability exist, none offer a standardized specification that is supported by all major hardware vendors. Additionally, experiences show that obtaining a high percentage of peak performance often requires architecture-specific code. This work details a pragmatic multi-architecture computational fluid dynamics library focused on aerospace problems across the speed range from low subsonic to hypersonic flows involving thermochemical nonequilibrium. A thin abstraction layer above NVIDIA CUDA C++ is utilized, which enables primarily single-source software currently capable of running efficiently on multicore CPUs, NVIDIA GPUs, AMD GPUs, and Intel GPUs. Results on various problems of interest across the speed range are presented and performance is compared between various architectures.

GPU↗

A Multi-Architecture Approach for Implicit Computational Fluid Dynamics on Unstructured Grids

High-performance computing (HPC) architectures are trending toward manycore paradigms such as graphics processing units (GPUs). Approximately half of the top 100 publicly disclosed supercomputers in the world utilize GPU accelerators for performance. This is in contrast to a decade ago, where there were only a few such machines in the top 100. It is not currently possible to compile and run legacy central processing unit (CPU) software efficiently on GPUs without significant refactoring. Though a number of frameworks offering performance portability exist, none offer a standardized specification that is supported by all major hardware vendors. Additionally, experiences show that obtaining a high percentage of peak performance often requires architecture-specific code. This work details a pragmatic multi-architecture computational fluid dynamics library focused on aerospace problems across the speed range from low subsonic to hypersonic flows involving thermochemical nonequilibrium. A thin abstraction layer above NVIDIA CUDA C++ is utilized, which enables primarily single-source software currently capable of running efficiently on multicore CPUs, NVIDIA GPUs, AMD GPUs, and Intel GPUs. Results on various problems of interest across the speed range are presented and performance is compared between various architectures.

GPU↗

An On Board Processor (OBP) for OAO C

A stored program computer and its application on OAO is considered. The parallel computer has a memory capacity of 16,384 words of 18 bits each, one central processor unit, two 4096 word memory units, and one input/output unit. The I/O has no direct data connection with the CPU so that all data flow between these two units must pass through memory by way of the memory data bus. The primary functions of the onboard computer are auxiliary command storage, spacecraft monitoring and malfunction reporting, data compression and status summary, and possible performance of emergency corrective action.

Hartenstein, R. G.↗

A workload model and measures for computer performance evaluation

A generalized workload definition is presented which constructs measurable workloads of unit size from workload elements, called elementary processes. An elementary process makes almost exclusive use of one of the processors, CPU, I/O processor, etc., and is measured by the cost of its execution. Various kinds of user programs can be simulated by quantitative composition of elementary processes into a type. The character of the type is defined by the weights of its elementary processes and its structure by the amount and sequence of transitions between its elementary processes. A set of types is batched to a mix. Mixes of identical cost are considered as equivalent amounts of workload. These formalized descriptions of workloads allow investigators to compare the results of different studies quantitatively. Since workloads of different composition are assigned a unit of cost, these descriptions enable determination of cost effectiveness of different workloads on a machine. Subsequently performance parameters such as throughput rate, gain factor, internal and external delay factors are defined and used to demonstrate the effects of various workload attributes on the performance of a selected large scale computer system.

Kerner, H.↗

A programmable computer interface for CAMAC

An interface has been developed for CAMAC instrumentation systems that implements data transfers controlled either by the computer CPU or by an autonomous (data-channel) processor in the interface unit. The data channel processor executes programs stored in the computer memory. These programs consist of standard CAMAC module commands plus special control characters and commands for the processor itself. The interface was built for the PDP-15 computer, which has an 18-bit word structure, but both 18- and 24-bit data transfers can be made. A software system has been written that exploits the many features of the processor.

Bercaw, R. W.↗

Digital image processing for the earth resources technology satellite data.

This paper discusses the problems of digital processing of the large volumes of multispectral image data that are expected to be received from the ERTS program. Correction of geometric and radiometric distortions are discussed and a byte oriented implementation is proposed. CPU timing estimates are given for a System/360 Model 67, and show that a processing throughput of 1000 image sets per week is feasible.

Will, P. M.↗

Study of efficient video compression algorithms for space shuttle applications

Results are presented of a study on video data compression techniques applicable to space flight communication. This study is directed towards monochrome (black and white) picture communication with special emphasis on feasibility of hardware implementation. The primary factors for such a communication system in space flight application are: picture quality, system reliability, power comsumption, and hardware weight. In terms of hardware implementation, these are directly related to hardware complexity, effectiveness of the hardware algorithm, immunity of the source code to channel noise, and data transmission rate (or transmission bandwidth). A system is recommended, and its hardware requirement summarized. Simulations of the study were performed on the improved LIM video controller which is computer-controlled by the META-4 CPU.

Poo, Z.↗

Space shuttle orbital maneuvering system failure detection and identification software requirements (uncontrolled)

Candidate designs and their software implementation are presented for the Orbital Maneuvering System (OMS) Failure Detection and Identification (FDI) algorithms in the Redundance Management (RM) module of the Space Shuttle Guidance, Navigation, and Control (GN&C) software. The OMS engine FDI algorithm monitors OMS engine thrust performance, and the OMS actuator FDI algorithm monitors OMS gimbal actuator performance. The software functional requirements of the algorithms are described along with the objective of each algorithm. A list of the assumptions which have governed its design, input/output requirements, a functional description of the algorithm (including a functional block diagram), and input interface requirements are given. The HAL (the language of the space shuttle flight computer) software formulation of the algorithms is considered including structured flowcharts of the procedures, estimates of flight computer core storage and CPU time, and processing requirements. A glossary of the symbols used to define the software requirements and formulations is included.

Damario, L. A.↗

A numerical comparison of discrete Kalman filtering algorithms: An orbit determination case study

The numerical stability and accuracy of various Kalman filter algorithms are thoroughly studied. Numerical results and conclusions are based on a realistic planetary approach orbit determination study. The case study results of this report highlight the numerical instability of the conventional and stabilized Kalman algorithms. Numerical errors associated with these algorithms can be so large as to obscure important mismodeling effects and thus give misleading estimates of filter accuracy. The positive result of this study is that the Bierman-Thornton U-D covariance factorization algorithm is computationally efficient, with CPU costs that differ negligibly from the conventional Kalman costs. In addition, accuracy of the U-D filter using single-precision arithmetic consistently matches the double-precision reference results. Numerical stability of the U-D filter is further demonstrated by its insensitivity of variations in the a priori statistics.

Thornton, C. L.↗

Evaluation of electrostatic charge effects on the data processing system and the orbiter communication and tracking receivers

An analysis of radiated interference test results obtained from frictionally charged Orbiter TPS tile was presented. The tests included the measurement of noise pick-up by Orbiter S-band, L-band, C-band, and Ku-band antennas located beneath the tiles in a manner simulating their installation on Orbiter. In addition, the radiated field characteristics resulting from the static discharge was determined. The results are analyzed as to their effect on data bus equipment and on Orbiter Communications and Tracking (C&T) receivers. It was concluded that the radiated interference should have no effect on MDM's. However the CPU, IOP and PMU enclosures require some minor modification to assure immunity from P-static interference. Orbiter antenna tests indicate that the S-band receiver should not be affected by P-static noise. The TACAN and Radar Altimeter performance appears to be adequate but with a small margin. MSBLS performance is uncertain because laboratory instrumentation cannot approach the MSBLS sensitivity.

Lawton, R. M.↗

System applications of the fault tolerant memory

Conventional memory technologies currently employed in aerospace applications contribute at least fifty percent to system unreliability (where the system includes CPU, I/O and memory). A fault tolerant memory performs both error correction and memory replacement at the bit plane level. To determine the effects of system design of using a fault tolerant memory in space applications, analysis was performed to determine tradeable hardware configurations that meet the reliability goals of each program. The candidate configurations, which included redundant elements of the computer system with both conventional and fault tolerant memories, were then traded in terms of selection criteria of cost, weight, volume, and power. These trade studies demonstrated that a fault tolerant memory provided significant advantages in terms of cost, weight, and volume. The memory selected for this analysis was a recently developed five fault tolerant memory.

Murphy, L. J.↗

Canonical analysis for increased classification speed and channel selection

The quadratic form can be expressed as a monotonically increasing sum of squares when the inverse covariance matrix is represented in canonical form. This formulation has the advantage that, in testing a particular class hypothesis, computations can be discontinued when the partial sum exceeds the smallest value obtained for other classes already tested. A method for channel selection is presented which arranges the original input measurements in that order which minimizes the expected number of computations. The classification algorithm was tested on data from LARS Flight Line C1 and found to reduce the sum-of-products operations by a factor of 6.7 in comparison with the conventional approach. In effect, the accuracy of a twelve-channel classification was achieved using only that CPU time required for a conventional four-channel classification.

Eppler, W.↗

A general method for calculating three-dimensional compressible laminar and turbulent boundary layers on arbitrary wings

The method described utilizes a nonorthogonal coordinate system for boundary-layer calculations. It includes a geometry program that represents the wing analytically, and a velocity program that computes the external velocity components from a given experimental pressure distribution when the external velocity distribution is not computed theoretically. The boundary layer method is general, however, and can also be used for an external velocity distribution computed theoretically. Several test cases were computed by this method and the results were checked with other numerical calculations and with experiments when available. A typical computation time (CPU) on an IBM 370/165 computer for one surface of a wing which roughly consist of 30 spanwise stations and 25 streamwise stations, with 30 points across the boundary layer is less than 30 seconds for an incompressible flow and a little more for a compressible flow.

Cebeci, T.↗

Output feedback regulator design for jet engine control systems

A multivariable control design procedure based on the output feedback regulator formulation is described and applied to turbofan engine model. Full order model dynamics, were incorporated in the example design. The effect of actuator dynamics on closed loop performance was investigaged. Also, the importance of turbine inlet temperature as an element of the dynamic feedback was studied. Step responses were given to indicate the improvement in system performance with this control. Calculation times for all experiments are given in CPU seconds for comparison purposes.

Merrill, W. C.↗

An efficient numerical method for solving the incompressible Navier-Stokes equations

This paper describes an efficient numerical method for solving the steady incompressible Navier-Stokes equations. The method is a fully implicit method based on the generalized Galerkin method, and the resulting system of equations is solved in a sweeping mode by iterative line relaxation. Results of the present method are compared with published results for separating and reattaching flows, and parametric studies showing the effects of step size, boundary locations, and Reynolds number are presented. The present method is substantially faster than previously published methods; typical run times range from 10 sec to 1 min of 7600 CPU time; and, based on results obtained to date, it is stable at any Reynolds number.

Murphy, J. D.↗

Numerical comparison of Kalman filter algorithms - Orbit determination case study

Numerical characteristics of various Kalman filter algorithms are illustrated with a realistic orbit determination study. The case study of this paper highlights the numerical deficiencies of the conventional and stabilized Kalman algorithms. Computational errors associated with these algorithms are found to be so large as to obscure important mismodeling effects and thus cause misleading estimates of filter accuracy. The positive result of this study is that the U-D covariance factorization algorithm has excellent numerical properties and is computationally efficient, having CPU costs that differ negligibly from the conventional Kalman costs. Accuracies of the U-D filter using single precision arithmetic consistently match the double precision reference results. Numerical stability of the U-D filter is further demonstrated by its insensitivity to variations in the a priori statistics.

Bierman, G. J.↗

Numerical comparison of discrete Kalman filter algorithms - Orbit determination case study

Numerical characteristics of various Kalman filter algorithms are illustrated with a realistic orbit determination study. The case study of this paper highlights the numerical deficiencies of the conventional and stabilized Kalman algorithms. Computational errors associated with these algorithms are found to be so large as to obscure important mismodeling effects and thus cause misleading estimates of filter accuracy. The positive result of this study is that the U-D covariance factorization algorithm has excellent numerical properties and is computationally efficient, having CPU costs that differ negligibly from the conventional Kalman costs. Accuracies of the U-D filter using single precision arithmetic consistently match the double precision reference results. Numerical stability of the U-D filter is further demonstrated by its insensitivity to variations in the a priori statistics.

Bierman, G. J.↗

Multipurpose system simulator

Multipurpose System Simulator (MPSS) evaluates relative performance of competitive computer systems and isolates areas for enhancement in existing or proposed systems. Model can simulate multiple central-processing-unit (CPU) interactive systems.

Packard, C. A.↗