Search NASA⌕ Search

SEARCH · Search NASA

Results for “Multiple processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Impacts of the IBM Cell Processor to Support Climate Models

NASA is interested in the performance and cost benefits for adapting its applications to the IBM Cell processor. However, its 256KB local memory per SPE and the new communication mechanism, make it very challenging to port an application. We selected the solar radiation component of the NASA GEOS-5 climate model, which: (1) is representative of column physics (approximately 50% computational time), (2) has a high computational load relative to transferring data from and to main memory, (3) performs independent calculations across multiple columns. We converted the baseline code (single-precision, Fortran) to C and ported it with manually SIMDizing 4 independent columns and found that a Cell with 8 SPEs can process 2274 columns per second. Compared with the baseline results, the Cell is approximately 5.2X, approximately 8.2X, approximately 15.1X faster than a core on Intel Woodcrest, Dempsey, and Itanium2, respectively. We believe this dramatic performance improvement makes a hybrid cluster with Cell and traditional nodes competitive.

Zhou, Shujia↗

Parallelization of irregularly coupled regular meshes

Regular meshes are frequently used for modeling physical phenomena on both serial and parallel computers. One advantage of regular meshes is that efficient discretization schemes can be implemented in a straight forward manner. However, geometrically-complex objects, such as aircraft, cannot be easily described using a single regular mesh. Multiple interacting regular meshes are frequently used to describe complex geometries. Each mesh models a subregion of the physical domain. The meshes, or subdomains, can be processed in parallel, with periodic updates carried out to move information between the coupled meshes. In many cases, there are a relatively small number (one to a few dozen) subdomains, so that each subdomain may also be partitioned among several processors. We outline a composite run-time/compile-time approach for supporting these problems efficiently on distributed-memory machines. These methods are described in the context of a multiblock fluid dynamics problem developed at LaRC.

Chase, Craig↗

Performance Analysis of a Hybrid Overset Multi-Block Application on Multiple Architectures

This paper presents a detailed performance analysis of a multi-block overset grid compu- tational fluid dynamics app!ication on multiple state-of-the-art computer architectures. The application is implemented using a hybrid MPI+OpenMP programming paradigm that exploits both coarse and fine-grain parallelism; the former via MPI message passing and the latter via OpenMP directives. The hybrid model also extends the applicability of multi-block programs to large clusters of SNIP nodes by overcoming the restriction that the number of processors be less than the number of grid blocks. A key kernel of the application, namely the LU-SGS linear solver, had to be modified to enhance the performance of the hybrid approach on the target machines. Investigations were conducted on cacheless Cray SX6 vector processors, cache-based IBM Power3 and Power4 architectures, and single system image SGI Origin3000 platforms. Overall results for complex vortex dynamics simulations demonstrate that the SX6 achieves the highest performance and outperforms the RISC-based architectures; however, the best scaling performance was achieved on the Power3.

Djomehri, M. Jahed↗

Effects of error sources on the parallelism of an optical matrix-vector processor

The error sources in a high accuracy optical matrix-vector processor are analyzed by numerical simulation in terms of their effects on the parallelism and speed of the processor. These effects are detailed for radices -2, -4 and -8. Radix -4 is shown to provide maximum parallel processing capabilities under the effects of the system's error sources. Processing speed is shown to be a function of matrix partitioning and the number of parallel processing channels. Consequently, radix -4 operation provides a higher processing speed than radix -2 and -8 for most matrix-vector multiplications when error source effects are considered.

Perlee, Caroline J.↗

Manned maneuvering unit applications for automated rendezvous and capture

Automated Rendezvous and Capture (AR&C) is an important technology to multiple National Aeronautics and Space Administration (NASA) programs and centers. The recent Johnson Spacecraft Center (JSC) AR&C Quality Function Deployment (QFD) has listed on-orbit demonstration of related technologies as a near term priority. Martin Marietta has been evaluating use of the Manned Maneuvering Unit (MMU) for a low cost near term on-orbit demonstration of AR&C technologies such as control algorithms, sensors, and processors as well as system level performance. The MMU Program began in 1979 as the method of repairing the Space Shuttle (STS) Thermal Protection System (the tiles). The units were not needed for this task, but were successfully employed during three Shuttle flights in 1984: a test flight was flown in in February as proof of concept, in April the MMU participated in the Solar Max Repair Mission, and in November the MMU's returned to space to successfully rescue the two errant satellites, Westar and Palapa. In the intervening years, the MMU simulator and MMU Qualification Test Unit (QTU) have been used for Astronaut training and experimental evaluations. The Extra-Vehicular Activities (EVA) Retriever has used the QTU, in an unmanned form, as a free-flyer on the Johnson Space Center (JSC) Precision Air Bearing Floor (PABF). Currently, the MMU is undergoing recertification for flight. The two flight units were removed from storage in September, 1991 and evaluation tests were performed. The tests demonstrated that the units are in good shape with no discrepancies that would preclude further use. The Return to Flight effort is currently clearing up recertification issues and evaluating the design against the present Shuttle environments.

Brehm, Donald L.↗

Implementation of Adaptive Digital Controllers on Programmable Logic Devices

Much has been made of the capabilities of FPGA's (Field Programmable Gate Arrays) in the hardware implementation of fast digital signal processing (DSP) functions. Such capability also makes and FPGA a suitable platform for the digital implementation of closed loop controllers. There are myriad advantages to utilizing an FPGA for discrete-time control functions which include the capability for reconfiguration when SRAM- based FPGA's are employed, fast parallel implementation of multiple control loops and implementations that can meet space level radiation tolerance in a compact form-factor. Other researchers have presented the notion that a second order digital filter with proportional-integral-derivative (PID) control functionality can be implemented in an FPGA. At Marshall Space Flight Center, the Control Electronics Group has been studying adaptive discrete-time control of motor driven actuator systems using digital signal processor (DSF) devices. Our goal is to create a fully digital, flight ready controller design that utilizes an FPGA for implementation of signal conditioning for control feedback signals, generation of commands to the controlled system, and hardware insertion of adaptive control algorithm approaches. While small form factor, commercial DSP devices are now available with event capture, data conversion, pulse width modulated outputs and communication peripherals, these devices are not currently available in designs and packages which meet space level radiation requirements. Meeting our goals requires alternative compact implementation of such functionality to withstand the harsh environment encountered on spacecraft. Radiation tolerant FPGA's are a feasible option for reaching these goals.

Gwaltney, David A.↗

Method and apparatus for pulse width modulation control of an AC induction motor

An inverter is connected between a source of DC power and a three-phase AC induction motor, and a micro-processor-based circuit controls the inverter using pulse width modulation techniques. In the disclosed method of pulse width modulation, both edges of each pulse of a carrier pulse train are equally modulated by a time proportional to sin .THETA., where .THETA. is the angular displacement of the pulse center at the motor stator frequency from a fixed reference point on the carrier waveform. The carrier waveform frequency is a multiple of the motor stator frequency. The modulated pulse train is then applied to each of the motor phase inputs with respective phase shifts of 120.degree. at the stator frequency. Switching control commands of electronic switches in the inverter are stored in a random access memory (RAM) and the locations of the RAM are successively read out in a cyclic manner, each bit of a given RAM location controlling a respective phase input of the motor. The DC power source preferably comprises rechargeable batteries and all but one of the electronic switches in the inverter can be disabled, the remaining electronic switch being part of a flyback DC-DC converter circuit for recharging the battery.

Geppert, Steven↗

Evaluation of the Monotonic Lagrangian Grid and Lat-Long Grid for Air Traffic Management

The Air Traffic Monotonic Lagrangian Grid (ATMLG) is used to simulate a 24 hour period of air traffic flow in the National Airspace System (NAS). During this time period, there are 41,594 flights over the United States, and the flight plan information (departure and arrival airports and times, and waypoints along the way) are obtained from an Federal Aviation Administration (FAA) Enhanced Traffic Management System (ETMS) dataset. Two simulation procedures are tested and compared: one based on the Monotonic Lagrangian Grid (MLG), and the other based on the stationary Latitude-Longitude (Lat- Long) grid. Simulating one full day of air traffic over the United States required the following amounts of CPU time on a single processor of an SGI Altix: 88 s for the MLG method, and 163 s for the Lat-Long grid method. We present a discussion of the amount of CPU time required for each of the simulation processes (updating aircraft trajectories, sorting, conflict detection and resolution, etc.), and show that the main advantage of the MLG method is that it is a general sorting algorithm that can sort on multiple properties. We discuss how many MLG neighbors must be considered in the separation assurance procedure in order to ensure a five-mile separation buffer between aircraft, and we investigate the effect of removing waypoints from aircraft trajectories. When aircraft choose their own trajectory, there are more flights with shorter duration times and fewer CD&R maneuvers, resulting in significant fuel savings.

Kaplan, Carolyn↗

Accelerating science: The usage of commercial clouds in ATLAS Distributed Computing

The ATLAS experiment at CERN is one of the largest scientific machines built to date and will have ever growing computing needs as the Large Hadron Collider collects an increasingly larger volume of data over the next 20 years. ATLAS is conducting R&D projects on Amazon Web Services and Google Cloud as complementary resources for distributed computing, focusing on some of the key features of commercial clouds: lightweight operation, elasticity and availability of multiple chip architectures. The proof of concept phases have concluded with the cloud-native, vendoragnostic integration with the experiment’s data and workload management frameworks. Google Cloud has been used to evaluate elastic batch computing, ramping up ephemeral clusters of up to O(100k) cores to process tasks requiring quick turnaround. Amazon Web Services has been exploited for the successful physics validation of the Athena simulation software on ARM processors. We have also set up an interactive facility for physics analysis allowing endusers to spin up private, on-demand clusters for parallel computing with up to 4 000 cores, or run GPU enabled notebooks and jobs for machine learning applications. The success of the proof of concept phases has led to the extension of the Google Cloud project, where ATLAS will study the total cost of ownership of a production cloud site during 15 months with 10k cores on average, fully integrated with distributed grid computing resources and continue the R&D projects.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Grid generation and inviscid flow computation about cranked-winged airplane geometries

An algebraic grid generation procedure that defines a patched multiple-block grid system suitable for fighter-type aircraft geometries with fuselage and engine inlet, canard or horizontal tail, cranked delta wing and vertical fin has been developed. The grid generation is based on transfinite interpolation and requires little computational power. A finite-volume Euler solver using explicit Runge-Kutta time-stepping has been adapted to this grid system and implemented on the VPS-32 vector processor with a high degree of vectorization. Grids are presented for an experimental aircraft with fuselage, canard, 70-20-cranked wing, and vertical fin. Computed inviscid compressible flow solutions are presented for Mach 2 at 3.79, 7 and 10 deg angles of attack. Conmparisons of the 3.79 deg computed solutions are made with available full-potential flow and Euler flow solutions on the same configuration but with another grid system. The occurrence of an unsteady solution in the 10 deg angle of attack case is discussed.

Eriksson, L.-E.↗

Using a Cray Y-MP as an array processor for a RISC Workstation

As microprocessors increase in power, the economics of centralized computing has changed dramatically. At the beginning of the 1980's, mainframes and super computers were often considered to be cost-effective machines for scalar computing. Today, microprocessor-based RISC (reduced-instruction-set computer) systems have displaced many uses of mainframes and supercomputers. Supercomputers are still cost competitive when processing jobs that require both large memory size and high memory bandwidth. One such application is array processing. Certain numerical operations are appropriate to use in a Remote Procedure Call (RPC)-based environment. Matrix multiplication is an example of an operation that can have a sufficient number of arithmetic operations to amortize the cost of an RPC call. An experiment which demonstrates that matrix multiplication can be executed remotely on a large system to speed the execution over that experienced on a workstation is described.

Lamaster, Hugh↗

Flash memory management system and method utilizing multiple block list windows

The present invention provides a flash memory management system and method with increased performance. The flash memory management system provides the ability to efficiently manage and allocate flash memory use in a way that improves reliability and longevity, while maintaining good performance levels. The flash memory management system includes a free block mechanism, a disk maintenance mechanism, and a bad block detection mechanism. The free block mechanism provides efficient sorting of free blocks to facilitate selecting low use blocks for writing. The disk maintenance mechanism provides for the ability to efficiently clean flash memory blocks during processor idle times. The bad block detection mechanism provides the ability to better detect when a block of flash memory is likely to go bad. The flash status mechanism stores information in fast access memory that describes the content and status of the data in the flash disk. The new bank detection mechanism provides the ability to automatically detect when new banks of flash memory are added to the system. Together, these mechanisms provide a flash memory management system that can improve the operational efficiency of systems that utilize flash memory.

Chow, James↗

Datascope to Enable Earth Independent Medical Operations (EIMO)

BACKGROUND: NASA has amassed sixty years of knowledge and experience relevant to maintenance of crew health and performance in low earth orbit. The Apollo Program introduced the importance of ensuring progressively autonomous operational capability. Earth Independent Medical Operations (EIMO) will require a gradual shift in the balance of medical responsibility, management, and authority from terrestrial to space-based assets. Terrestrial assets will continue to be essential for pre-mission screening and planning in addition to maintenance of crew health and performance. However, new capabilities are needed to enable EIMO and the amount of data required to support these systems, and mitigate the impacts of data transmission delays and reduced bandwidth coupled with lack of cloud-like resources and on-board computing capacity that is currently unclear or operationally insufficient. OVERVIEW: The overall goal of EIMO is to develop artificial intelligence (AI)-based solutions to analyze crew health and performance data utilizing a clinical decision support system (CDSS) to provide crew medical officers (CMO) with the equivalent of real-time, on-board medical consults. The EIMO ecosystem is envisioned as a “system of systems” where embedded reference databases and real-time data streams from multiple input vectors continuously and seamlessly assess crew health and performance. EIMO will be designed to make recommendations to the CMO using multi-modal AI-based natural language processing and machine learning methods with interoperability to push/pull data within and between multiple vehicle and habitat architectures. DISCUSSION: Data flows and storage/retrieval capacity are severely constrained during space missions and the challenges will become even greater during exploration missions. Just as each past program from Mercury to the International Space Station (ISS) required rethinking the interaction between ground-based controllers and space-based crew, so too will future missions to the Moon and Mars. While the NASA High-Performance Spaceflight Computing Processor project aims to increase computational capacity by 100 times over current spaceflight computers, the projected deliverable still lags considerably behind what will be needed to enable an AI-driven CDSS. Restrictions in processing speed and data storage capacity, coupled with transmission bottlenecks and delays, necessitate definition and optimization of an integrated data architecture to enable a progressively autonomous medical capability.

Medical operations↗

Precise micromotion compensation of a tilted ion chain

Excess micromotion can be a substantial source of errors in trapped-ion based quantum processors and clocks due to the sensitivity of the internal states of the ion to external fields and motion. This problem can be fixed by compensating background electric fields in order to position ions at the RF node and minimize their driven micromotion. Here we describe techniques for compensating ion chains in scalable surface ion traps. These traps are capable of cancelling stray electric fields with fine spatial resolution in order to compensate multiple closely spaced ions due to their large number of relatively small control electrodes. We demonstrate a technique that compensates an ion chain to better than 5 V/m and within 0.1 degrees of chain rotation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Monolithic Microwave Integrated Circuit (MMIC) technology for space communications applications

Future communications satellites are likely to use gallium arsenide (GaAs) monolithic microwave integrated-circuit (MMIC) technology in most, if not all, communications payload subsystems. Multiple-scanning-beam antenna systems are expected to use GaAs MMICs to increase functional capability, to reduce volume, weight, and cost, and to greatly improve system reliability. RF and IF matrix switch technology based on GaAs MMICs is also being developed for these reasons. MMIC technology, including gigabit-rate GaAs digital integrated circuits, offers substantial advantages in power consumption and weight over silicon technologies for high-throughput, on-board baseband processor systems. For the more distant future pseudomorphic indium gallium arsenide (InGaAs) and other advanced III-V materials offer the possibility of MMIC subsystems well up into the millimeter wavelength region. All of these technology elements are in NASA's MMIC program. Their status is reviewed.

Connolly, Denis J.↗

Advances in gallium arsenide monolithic microwave integrated-circuit technology for space communications systems

Future communications satellites are likely to use gallium arsenide (GaAs) monolithic microwave integrated-circuit (MMIC) technology in most, if not all, communications payload subsystems. Multiple-scanning-beam antenna systems are expected to use GaAs MMIC's to increase functional capability, to reduce volume, weight, and cost, and to greatly improve system reliability. RF and IF matrix switch technology based on GaAs MMIC's is also being developed for these reasons. MMIC technology, including gigabit-rate GaAs digital integrated circuits, offers substantial advantages in power consumption and weight over silicon technologies for high-throughput, on-board baseband processor systems. In this paper, current developments in GaAs MMIC technology are described, and the status and prospects of the technology are assessed.

Bhasin, K. B.↗

Monolithic Microwave Integrated Circuit (MMIC) technology for space communications applications

Future communications satellites are likely to use gallium arsenide (GaAs) monolithic microwave integrated-circuit (MMIC) technology in most, if not all, communications payload subsystems. Multiple-scanning-beam antenna systems are expected to use GaAs MMIC's to increase functional capability, to reduce volume, weight, and cost, and to greatly improve system reliability. RF and IF matrix switch technology based on GaAs MMIC's is also being developed for these reasons. MMIC technology, including gigabit-rate GaAs digital integrated circuits, offers substantial advantages in power consumption and weight over silicon technologies for high-throughput, on-board baseband processor systems. For the more distant future pseudomorphic indium gallium arsenide (InGaAs) and other advanced III-V materials offer the possibility of MMIC subsystems well up into the millimeter wavelength region. All of these technology elements are in NASA's MMIC program. Their status is reviewed.

Connolly, Denis J.↗

Multichannel Phase and Power Detector

An electronic signal-processing system determines the phases of input signals arriving in multiple channels, relative to the phase of a reference signal with which the input signals are known to be coherent in both phase and frequency. The system also gives an estimate of the power levels of the input signals. A prototype of the system has four input channels that handle signals at a frequency of 9.5 MHz, but the basic principles of design and operation are extensible to other signal frequencies and greater numbers of channels. The prototype system consists mostly of three parts: An analog-to-digital-converter (ADC) board, which coherently digitizes the input signals in synchronism with the reference signal and performs some simple processing; A digital signal processor (DSP) in the form of a field-programmable gate array (FPGA) board, which performs most of the phase- and power-measurement computations on the digital samples generated by the ADC board; and A carrier board, which allows a personal computer to retrieve the phase and power data. The DSP contains four independent phase-only tracking loops, each of which tracks the phase of one of the preprocessed input signals relative to that of the reference signal (see figure). The phase values computed by these loops are averaged over intervals, the length of which is chosen to obtain output from the DSP at a desired rate. In addition, a simple sum of squares is computed for each channel as an estimate of the power of the signal in that channel. The relative phases and the power level estimates computed by the DSP could be used for diverse purposes in different settings. For example, if the input signals come from different elements of a phased-array antenna, the phases could be used as indications of the direction of arrival of a received signal and/or as feedback for electronic or mechanical beam steering. The power levels could be used as feedback for automatic gain control in preprocessing of incoming signals. For another example, the system could be used to measure the phases and power levels of outputs of multiple power amplifiers to enable adjustment of the amplifiers for optimal power combining.

Li, Samuel↗