Single-event upset in power-PC processors
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The objective of this coarse Single Event Effect (SEE) test is to determine the suitability of the commercial Virtex-II Pro family for use in spaceflight applications. To this end, this test is primarily intended to determine any Singe Event Latchup (SEL) susceptibilities for these devices. Secondly, this test is intended to measure the level of Single Event Upset (SEU) susceptibilities and in a general sense where they occur. The coarse SEE test was performed on a commercial XC2VP7 device, a relatively small single processor version of the Virtex-II Pro. As the XC2VP7 shares the same functional block design and fabrication process with the larger Virtex-II Pro devices, the results of this test should also be applicable to the larger devices. The XC2VP7 device was tested on a commercial Virtex-II Pro development board. The testing was performed at the Cyclotron laboratories at Texas A&M and Michigan State Universities using ions of varying energy levels and fluences.
Motivation for this work is: (1) Accurately characterize digital signal processor (DSP) core single-event effect (SEE) behavior (2) Test DSP cores across a large frequency range and across various input conditions (3) Isolate SEE analysis to DSP cores alone (4) Interpret SEE analysis in terms of single-event upsets (SEUs) and single-event transients (SETs) (5) Provide flight missions with accurate estimate of DSP core error rates and error signatures.
The Xilinx Vix-11 Pro is a platform FPGA that embeds multiple microprocessors within the fabric of an SRAM-based reprogrammable FPGA. The variety and quantity of resources provided by this family of devices make them very attractive for spaceflight applications. However,these devices will be susceptible to single event effects (SEE), which must be mitigated. Observations from prior testing of the Xilinx Virtex-II Pro suggest that the PowerPC core has significant vulnerability to SEES. However, these initial tests were not designed to exclusively target the functionality of the PowerPC, therefore making it difficult to distinguish processor upsets from fabric upsets. The main focus of this paper involves detailed SEE testing of the embedded PowerPC core. Due to the complexity of the PowerPC, various custom test applications, both static and dynamic, will be designed to isolate each Unit of the processor. Collective analysis of the test results will provide insight into the exact upset mechanism of the PowerPC. With this information, mitigations schemes can be developed and tested that address the specific susceptibilities of these devices. The test bed will be the Xilinx SEE Consortium Virtex-II Pro test board, which allows for configuration scrubbing, design triplication, and ease of data collection. Testing will be performed at the Indiana University Cyclotron Facility using protons of varying energy levels and fluencies. This paper will present the detailed test approach along with the results.
The Xilinx Virtex-II Pro is a platform FPGA that embeds multiple microprocessors within the fabric of an SRAM-based reprogrammable FPGA. The variety and quantity of resources provided by this family of devices make them very attractive for spaceflight applications. However, these devices will be susceptible to single event effects (SEE), which must be mitigated. To use the Virtex-II Pro reliably in space applications, these devices must first be tested to determine if they are susceptible to single event latchup (SEL), the degree to which they are susceptible to single event upsets (SEU) and single event transients (SET), and how these effects are manifested in the device. With this information, mitigations schemes can be developed and tested that address the specific susceptiblities of these devices. This initial SEE test uses a commercial off the shelf Virtex-II Pro evaluation board, with a single processor XC2VP7 FPGA. The FPGA on this board is an acid etched device, which can be partially covered with a shield. The shield covers a portion of the logic, routing, and memory resources along with some of the RocketIO transceivers. The processor, along with a large portion of logic, routing, memory, and transceivers are left exposed. This test will be performed at the Cyclotron Laboratories at Texas A&M University and Michigan State University using ions of varying energy levels and fluencies.
The Applied Physics Laboratory (APL) has developed a magnetometer instrument for a swedish satellite named Freja with launch scheduled for August 1992 on a Chinese Long March rocket. The magnetometer controller utilized a custom microprocessor designed at APL with the Genesil silicon compiler. The processor evolved from our experience with an older bit-slice design and two prior single chip efforts. The architecture of our microprocessor greatly lowered software development costs because it was optimized to provide an interactive and extensible programming environment hosted by the target hardware. Radiation tolerance of the microprocessor was also tested and was adequate for Freja's mission -- 20 kRad(Si) total dose and very infrequent latch-up and single event upset events.
The Reconfigurable Hardware in Orbit (RHinO) project is focused on creating a set of design tools that facilitate and automate design techniques for reconfigurable computing in space, using SRAM-based field-programmable-gate-array (FPGA) technology. These tools leverage an established FPGA design environment and focus primarily on space effects mitigation and power optimization. The project is creating software to automatically test and evaluate the single-event-upsets (SEUs) sensitivities of an FPGA design and insert mitigation techniques. Extensions into the tool suite will also allow evolvable algorithm techniques to reconfigure around single-event-latchup (SEL) events. In the power domain, tools are being created for dynamic power visualiization and optimization. Thus, this technology seeks to enable the use of Reconfigurable Hardware in Orbit, via an integrated design tool-suite aiming to reduce risk, cost, and design time of multimission reconfigurable space processors using SRAM-based FPGAs.
This paper examines single-event upsets in advanced commercial SOI microprocessors in a dynamic mode, studying SEU sensitivity of General Purpose Registers (GPRs) with clock frequency. Results are presented for SOI processors with feature sizes of 0.18 microns and two different core voltages. Single-event upset from heavy ions is measured for advanced commercial microprocessors in a dynamic mode with clock frequency up to 1GHz. Frequency and core voltage dependence of single-event upsets in registers is discussed.
SEU from heavy-ions is measured for SOI PowerPC microprocessors. Results for 0.13 micron PowerPC with 1.1V core voltages increases over 1.3V versions. This suggests that improvement in SEU for scaled devices may be reversed. In recent years there has been interest in the possible use of unhardened commercial microprocessors in space because of their superior performance compared to hardened processors. However, unhardened devices are susceptible to upset from radiation space. More information is needed on how they respond to radiation before they can be used in space. Only a limited number of advanced microprocessors have been subjected to radiation tests, which are designed with lower clock frequencies and higher internal core voltage voltages than recent devices [1-6]. However the trend for commercial Silicon-on-insulator (SOI) microprocessors is to reduce feature size and internal core voltage and increase the clock frequency. Commercial microprocessors with the PowerPC architecture are now available that use partially depleted SOI processes with feature size of 90 nm and internal core voltage as low as 1.0 V and clock frequency in the GHz range. Previously, we reported SEU measurements for SOI commercial PowerPCs with feature size of 0.18 and 0.13 m [7, 8]. The results showed an order of magnitude reduction in saturated cross section compared to CMOS bulk counterparts. This paper examines SEUs in advanced commercial SOI microprocessors, focusing on SEU sensitivity of D-Cache and hangs with feature size and internal core voltage. Results are presented for the Motorola SOI processor with feature sizes of 0.13 microns and internal core voltages of 1.3 and 1.1 V. These results are compared with results for the Motorola SOI processors with feature size of 0.18 microns and internal core voltage of 1.6 and 1.3 V.
This research describes the development of an experimental radiation testing environment to investigate the single event effect (SEE) susceptibility of the 486-DX4 microprocessor. SEE effects are caused by radiation particles that disrupt the logic state of an operating semiconductor, and include single event upsets (SEU) and single event latchup (SEL). The relevance of this work can be applied directly to digital devices that are used in spaceflight computer systems. The 486-DX4 is a powerful commercial microprocessor that is currently under consideration for use in several spaceflight systems. As part of its selection process, it must be rigorously tested to determine its overall reliability in the space environment, including its radiation susceptibility. The goal of this research is to experimentally test and characterize the single event effects of the 486-DX4 microprocessor using a cyclotron facility as the fault-injection source. The test philosophy is to focus on the "operational susceptibility," by executing real software and monitoring for errors while the device is under irradiation. This research encompasses both experimental and analytical techniques, and yields a characterization of the 486-DX4's behavior for different operating modes. Additionally, the test methodology can accommodate a wide range of digital devices, such as microprocessors, microcontrollers, ASICS, and memory modules, for future testing. The goals were achieved by testing with three heavy-ion species to provide different linear energy transfer rates, and a total of six microprocessor parts were tested from two different vendors. A consistent set of error modes were identified that indicate the manner in which the errors were detected in the processor. The upset cross-section curves were calculated for each error mode, and the SEU threshold and saturation levels were identified for each processor. Results show a distinct difference in the upset rate for different configurations of the on-chip cache, as well as proving that one vendor is superior to the other in terms of latchup susceptibility. Results from this testing were also used to provide a mean-time-between-failure estimate of the 486-DX4 operating in the radiation environment for the International Space Station.
In 1995 NASA began an experimental program to develop a reusable crew return vehicle (CRV) for the International Space Station. The purpose of the CRV was threefold: (i) to bring home an injured or ill crewmember; (ii) to bring home the entire crew if the Shuttle fleet was grounded; and (iii) to evacuate the crew in the case of an imminent Station threat (i.e., fire, decompression, etc). Built at the Johnson Space Center, were two approach and landing prototypes and one spacecraft demonstrator (called V201). A series of increasingly complex ground subsystem tests were completed, and eight successful high-altitude drop tests were achieved to prove the design concept. In this program, an unprecedented amount of commercial-off-the-shelf technology was utilized in this first crewed spacecraft NASA has built since the Shuttle program. Unfortunately, in 2002 the program was canceled due to changing Agency priorities. The vehicle was 80% complete and the program was shut down in such a manner as to preserve design, development, test and engineering data. This paper describes the X-38 V201 fault-tolerant avionics system. Based on Draper Laboratory's Byzantine-resilient fault-tolerant parallel processing system and their "network element" hardware, each flight computer exchanges information on a strict timescale to process input data, compare results, and issue voted vehicle output commands. Major accomplishments achieved in this development include: (i) a space qualified two-fault tolerant design using mostly COTS (hardware and operating system); (ii) a single event upset tolerant network element board, (iii) on-the-fly recovery of a failed processor; (iv) use of synched cache; (v) realignment of memory to bring back a failed channel; (vi) flight code automatically generated from the master measurement list; and (vii) built in-house by a team of civil servants and support contractors. This paper will present an overview of the avionics system and the hardware implementation, as well as the system software and vehicle command & telemetry functions. Potential improvements and lessons learned on this program are also discussed.
The growth in data rates of instruments on future NASA spacecraft continues to outstrip the improvement in communications bandwidth and processing capabilities of radiation-hardened computers. Sophisticated autonomous operations strategies will further increase the processing workload. Given the reductions in spacecraft size and available power, standard radiation hardened computing systems alone will not be able to address the requirements of future missions. The REE project was intended to overcome this obstacle by developing a COTS- based supercomputer suitable for use as a science and autonomy data processor in most space environments. This development required a detailed knowledge of system behavior in the presence of Single Event Effect (SEE) induced faults so that mitigation strategies could be designed to recover system level reliability while maintaining the COTS throughput advantage. The REE project has developed a suite of tools and a methodology for predicting SEU induced transient fault rates in a range of natural space environments from ground-based radiation testing of component parts. In this paper we provide an overview of this methodology and tool set with a concentration on the radiation fault model and its use in the REE system development methodology. Using test data reported elsewhere in this and other conferences, we predict upset rates for a particular COTS single board computer configuration in several space environments.
NASA/Marshall Space Flight Center (MSFC) is continually looking for ways to reduce the costs and schedule and minimize the technical risks during the development of microgravity programs. One of the more prominent ways to minimize the cost and schedule is to use off-the-shelf hardware (OTS). However, the use of OTS often increases the risk. This paper addresses relevant factors considered during the selection and utilization of commercial off-the-shelf (COTS) flight computer processing equipment for the control of space station microgravity experiments. The paper will also discuss how to minimize the technical risks when using COTS processing hardware. Two microgravity experiments for which the COTS processing equipment is being evaluated for are the Equiaxed Dendritic Solidification Experiment (EDSE) and the Self-diffusion in Liquid Elements (SDLE) experiment. Since MSFC is the lead center for Microgravity research, EDSE and SDLE processor selection will be closely watched by other experiments that are being designed to meet payload carrier requirements. This includes the payload carriers planned for the International Space Station (ISS). The purpose of EDSE is to continue to investigate microstructural evolution of, and thermal interactions between multiple dendrites growing under diffusion controlled conditions. The purpose of SDLE is to determine accurate self-diffusivity data as a function of temperature for liquid elements selected as representative of class-like structures. In 1999 MSFC initiated a Center Director's Discretionary Fund (CDDF) effort to investigate and determine the optimal commercial data bus architecture that could lead to faster, better, and lower cost data acquisition systems for the control of microgravity experiments. As part of this effort various commercial data acquisition systems were acquired and evaluated. This included equipment with various form factors, (3U, 6U, others) and equipment that utilized various bus structures, (VME, PC104, STD bus). This evaluation of hardware was performed in conjunction with a trade study that considered over twenty (20) different factors relevant to the selection of an optimum design approach. These factors included; safety, sizing and timing, radiation hardness and single event upset, power consumption, heat dissipation, size and volume, expected service life, maintainability, heritage, operating systems, requirements for software reuse, availability of compatible interface boards, relative cost, schedule, reliability, EMI/EMC factors, "hot swap" capability, standards for conduction cooling, I/O capabilities, unique carrier requirements and operating system considerations. The approach to evaluate Safety as part of this study included a review of the Preliminary Hazard Analysis (PHA) for each of the experiment designs and a determination of how each hazard could be addressed and eliminated when different processors were selected. This included evaluating various design approaches and trade-offs between fault tolerant designs and fail-safe designs in accordance with NSTS 1700.7B. This will include the results of radiation testing where available. Various operating systems, such as VxWorks, Linux, QNX, and Embedded NT are evaluated and the advantages and disadvantages of their utilization are also addressed. Design implementation strategies for the various operating systems are considered and discussed. This paper presents the results and recommendations from this trade study. Preliminary conclusions from this study are that safety concerns from lack or radiation testing on COTS equipment can be addressed by additional testing and design considerations, the PC104 bus provided adequate I/O for the SDLE and EDSE microgravity experiments, and PC104 bus components offered significant advantages over VME and cPCI for weight and space reductions.
The two main objectives of this trade study are characterizing the mission radiation environment for multiple Design Reference Missions (DRM) and analysis of radiation effects on avionics with the goal of producing radiation tolerant Neuromorphic Computing processor chips with innovative radiation-induced fault mitigation. The NASA process for defining radiation requirements for flight avionics is applied to this domain. The effects of trapped protons and electrons in the Van Allen radiation belts predominates in Low-Earth-Orbit (LEO) and the solar wind, solar flares and Galactic Cosmic Rays (GCRs) are the dominant radiation challenge in the open space between the planets of our solar system. The nature and energy of the particles that cause circuit upset and failure is very different in the two regimes. Two results are produced from the radiation models: determining the Total Integrated Dose (TID) experienced by avionics for a given DRM and predicting the Single Event Effects (SEE) rates for avionics during high rate exposure. These tools are applied to existing semiconductors and can be used for predicting the radiation performance of future semiconductors based on early radiation testing of new devices.
Published data on the processors sensitivites with respect to SEU is generally obtained from radiation ground testing during which the program is executed by the DUT consists in the sequential inspection of each of the processor memory cells accessible to the user, through the execution of a suitable instruction sequence. In such programs, so-called static tests, typically considered memory cells are general-purpose registers, special registers (program counter, stack pointer...) and internal memory. Nevertheless, the register use and duty cycle of the final application will be very different, including using instructions no in the static tests and disturbing other potential SEU targets. The ideal would be the use of the final application program for the radiation ground testing, but generally this program is either unknown or unavailable when the qualification testing is performed on candidate circuits to space projects.