Search NASA⌕ Search

SEARCH · Search NASA

Results for “hardware reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

A real time, FEM based optimal control algorithm and its implementation using parallel processing hardware (transistors) in a microprocessor environment

There is an evident need to discover a means of establishing reliable, implementable controls for systems that are plagued by nonlinear and, or uncertain, model dynamics. The development of a generic controller design tool for tough-to-control systems is reported. The method utilizes a moving grid, time infinite element based solution of the necessary conditions that describe an optimal controller for a system. The technique produces a discrete feedback controller. Real time laboratory experiments are now being conducted to demonstrate the viability of the method. The algorithm that results is being implemented in a microprocessor environment. Critical computational tasks are accomplished using a low cost, on-board, multiprocessor (INMOS T800 Transputers) and parallel processing. Progress to date validates the methodology presented. Applications of the technique to the control of highly flexible robotic appendages are suggested.

Patten, William Neff↗

Fault diagnostic instrumentation design for environmental control and life support systems

As a development phase moves toward flight hardware, the system availability becomes an important design aspect which requires high reliability and maintainability. As part of continous development efforts, a program to evaluate, design, and demonstrate advanced instrumentation fault diagnostics was successfully completed. Fault tolerance designs for reliability and other instrumenation capabilities to increase maintainability were evaluated and studied.

Yang, P. Y.↗

The US Space Station programme

The Manned Space Station (MSS) involves NASA, and other countries, in the operation, maintenance and expansion of a permanent space facility. The extensive use of automation and robotics will advance those fields, and experimentation will be carried out in scientific and potentially commercial projects. The MSS will provide a base for astronomical observations, spacecraft assembly, refurbishment and repair, transportation intersection, staging for interplanetary exploration, and storage. Finally, MSS operations will be performed semi-autonomously from ground control. Phase B analysis is nearing completion, and precedes hardware development. Studies are being performed on generic advanced technologies which can reliably and flexibly be incorporated into the MSS, such as attitude control and stabilization, power, thermal, environmental and life support control, auxiliary propulsion, data management, etc. Guidelines are also being formulated regarding the areas of participation by other nations.

Hodge, J. D.↗

Advanced life support control/monitor instrumentation concepts for flight application

Development of regenerative Environmental Control/Life Support Systems requires instrumentation characteristics which evolve with successive development phases. As the development phase moves toward flight hardware, the system availability becomes an important design aspect which requires high reliability and maintainability. This program was directed toward instrumentation designs which incorporate features compatible with anticipated flight requirements. The first task consisted of the design, fabrication and test of a Performance Diagnostic Unit. In interfacing with a subsystem's instrumentation, the Performance Diagnostic Unit is capable of determining faulty operation and components within a subsystem, perform on-line diagnostics of what maintenance is needed and accept historical status on subsystem performance as such information is retained in the memory of a subsystem's computerized controller. The second focus was development and demonstration of analog signal conditioning concepts which reduce the weight, power, volume, cost and maintenance and improve the reliability of this key assembly of advanced life support instrumentation. The approach was to develop a generic set of signal conditioning elements or cards which can be configured to fit various subsystems. Four generic sensor signal conditioning cards were identified as being required to handle more than 90 percent of the sensors encountered in life support systems. Under company funding, these were detail designed, built and successfully tested.

Heppner, D. B.↗

The University of Hawaii Institute for Astronomy CCD camera control system

The University of Hawaii Institute for Astronomy CCD Camera Control System consists of a NeXT workstation, a graphical user interface, and a fiber optics communications interface which is connected to a San Diego State University CCD controller. The UH system employs the NeXT-resident Motorola DSP 56001 as a real time hardware controller. The DSP 56001 is interfaced to the Mach-based UNIX of the NeXT workstation by DMA and multithreading. Since the SDSU controller also uses the DPS 56001, the NeXT is used as a development platform for the embedded control software. The fiber optic interface links the two DSP 56001's through their Synchronous Serial Interfaces. The user interface is based on the NeXTStep windowing system. It is easy to use and features real-time display of image data and control over all camera functions. Both Loral and Tektronix 2048 x 2048 CCD's have been driven at full readout speeds, and the system is intended to be capable of simultaneous readout of four such CCD's. The total hardware package is compact enough to be quite portable and has been used on five different telescopes on Mauna Kea. The complete CCD control system can be assembled for a very low cost. The hardware and software of the control system has proven to be quite reliable, well adapted to the needs of astronomers, and extensible to increasingly complicated control requirements.

Jim, K. T. C.↗

An Overview of Advanced Data Acquisition System (ADAS)

The paper discusses the following: 1. Historical background. 2. What is ADAS? 3. R and D status. 4. Reliability/cost examples (1, 2, and 3). 5. What's new? 6. Technical advantages. 7. NASA relevance. 8. NASA plans/options. 9. Remaining R and D. 10. Applications. 11. Product benefits. 11. Commercial advantages. 12. intellectual property. Aerospace industry requires highly reliable data acquisition systems. Traditional Acquisition systems employ end-to-end hardware and software redundancy. Typically, redundancy adds weight, cost, power consumption, and complexity.

Mata, Carlos T.↗

Advanced Health Management Algorithms for Crew Exploration Applications

Achieving the goals of the President's Vision for Exploration will require new and innovative ways to achieve reliability increases of key systems and sub-systems. The most prominent approach used in current systems is to maintain hardware redundancy. This imposes constraints to the system and utilizes weight that could be used for payload for extended lunar, Martian, or other deep space missions. A technique to improve reliability while reducing the system weight and constraints is through the use of an Advanced Health Management System (AHMS). This system contains diagnostic algorithms and decision logic to mitigate or minimize the impact of system anomalies on propulsion system performance throughout the powered flight regime. The purposes of the AHMS are to increase the probability of successfully placing the vehicle into the intended orbit (Earth, Lunar, or Martian escape trajectory), increase the probability of being able to safely execute an abort after it has developed anomalous performance during launch or ascent phases of the mission, and to minimize or mitigate anomalies during the cruise portion of the mission. This is accomplished by improving the knowledge of the state of the propulsion system operation at any given turbomachinery vibration protection logic and an overall system analysis algorithm that utilizes an underlying physical model and a wide array of engine system operational parameters to detect and mitigate predefined engine anomalies. These algorithms are generic enough to be utilized on any propulsion system yet can be easily tailored to each application by changing input data and engine specific parameters. The key to the advancement of such a system is the verification of the algorithms. These algorithms will be validated through the use of a database of nominal and anomalous performance from a large propulsion system where data exists for catastrophic and noncatastrophic propulsion sytem failures.

Davidson, Matt↗

Using Systems Engineering to Develop an Integrated Crew Health and Performance System to Mitigate Risk for Human Exploration Missions

New space exploration missions are currently being designed to take humanity beyond low Earth orbit (LEO) to cis-lunar space, the lunar surface, and eventually to Mars. These missions carry increased risks due to a number of factors, including distance from Earth, exposure to deep space hazards, reduced capacity and ability to resupply and evacuate, and increased communication delays. As distance from Earth grows and mission length increases, a growing proportion of mission risk can be attributed to the “human system.” Almost twenty years ago, the Institute of Medicine in the United States recommended the early and complete integration of the human system into the spacecraft and mission design process to mitigate this increased risk inherent in exploration missions. Exploration missions will require increasing levels of crew self-sufficiency that will challenge the current operational paradigms established in LEO. Enabling progressive Earth independence requires management of the increasingly complex interactions among spacecraft systems and integration of all the data and functions that affect human health and performance into one coordinated system –the Crew Health and Performance (CHP) system. This system is a critical spacecraft system that is analogous to other systems such as propulsion, guidance and navigation, or avionics. Design and integration of the CHP system requires the evidence-based merger of typically disparate disciplines such as medicine, human factors, and life support system design, using systems engineering (SE) practices. SE provides traceability of spacecraft requirements and enables reliable and repeatable trade space analysis of the many competing options for hardware and software to be included in the mission architecture. The approach described here enables spacecraft and mission designers to align the scope of a CHP system with mission specific requirements, decrease risk to the crew, and increase the probability of mission success.

Kerry McGuire↗

Genesis Solar Wind – Capture, Return, Curate and Analyze: Looking Backward and Creating a Timeline

Introduction: In 1997 NASA’S Discovery Program selected the Genesis mission proposal to return solar wind samples to Earth for laboratory analyses. Principal Investigator Donald S. Burnett and the science team defined the purity of collector materials and ability to analyze solar wind composition to the precision required for planetary science. As a small mission, focused on a well-defined science goal, yet needing careful attention to engineering details, the communication among scientists and engineers, nurtured by Don Burnett, was exceptional. Genesis Mission and Curation Legacy: Genesis, as the first U. S. spacecraft to return astromaterial samples since Apollo, not only integrated the mission planning and flight teams, but also the science and sample curation teams during the mission development period. Since Genesis is a sample return mission, the Science Team was essential in certifying the collectors (sample containers for solar atoms). From inception, Genesis established mission funding for returned sample curation. JSC was lead in contamination control during mission preparation, including establishment of an ISO 4 cleanroom facility and use of ultrapure water (UPW) for cleaning flight hardware (and, as it turned out, for cleaning collectors after the mishap). Reliable, fast communication among scientists, engineers and curators at the hands-on level established deep respect among team members and efficient decision-making. JSC’s 50-years of astromaterial sample curation provided experienced sample processors onsite during recovery in Utah (a deep bench for emergency response). Post-recovery curation included iterative collaboration with science sample users to clean or verify cleanliness of samples. The science legacy from Genesis is addressed by Burnett and Jurewicz, this volume. In The Beginning: After Apollo sample return, Burnett and Marcia Neugebauer at JPL began discussing a solar wind sample return, with Neugebauer arguing that separate collection of solar wind regimes was essential science. By 1992 a solar wind sample return mission was presented at a workshop, and by 1994 a mission was proposed named Suess-Urey. The mission was re-proposed under a new name GENESIS and selected in 1997. Susan Niebur captured the Genesis mission history and stories, from high level management documents and from many interviews with participants [2]. Her account lets readers glimpse personality of participants in quotations from interviews. Need and Scope for Detailed Technical Timeline: A timeline constructed from lower level task documents has been initiated to document the resources and skills actually used, as well as task sequence or concurrency. Timelines for high level mission events are captured in two documents [1] [2] and for detailed re-entry events in [3]. A detailed technical timeline for Genesis mission and curation activities will provide data points for lower level tasks, such as ISO 4 curation facility construction time, preparation for nominal sample field recovery, mishap recovery, and UPW expansion. Changes in technology context 1990-2024: Semiconductor technologies were easily accessible in the U.S.A. (1990-1999), and the Genesis team used those resources for cleanroom design and UPW system expansion. Image documentation was changing from film to digital during cleanroom construction and payload cleaning (1997-2001). Engineering design was done using computer aided design proprietary software, making more difficult the archiving of payload configuration and materials. Email of documents, tracked delivery service and virtual meeting capability greatly improved communication efficiency. Information sources – Pre-launch mission preparation: Examples of mission science, engineering and contamination control are collector purity testing, payload design/fabrication and ISO 4 cleanroom construction. Information on timing of these activities comes from facility readiness reviews, management reviews, shipping documents, procurement documents, test reports, travel documents, laboratory logs, Quality Assurance documents, dates on images, participant notebooks and emails. Information sources – Sample return re-entry and field recovery activities: Information comes from event timelines produced by Mid-Air Recovery team, Lockheed team lead notes and from chase video, JPL Quality Assurance. Information also comes from images and logbooks from UTTR cleanroom operations and from curatorial documents. Information sources – Resulting science and sample cleaning processes: Agendas from the annual gatherings of the science team initially trace testing for collector purity/cleanliness, and after sample recovery, include collector cleaning and cleanliness assessment. Post-recovery documents include curatorial orders and procedures, sample allocation documents and LPSC abstracts. Timeline Objectives: A simple spreadsheet timeline with headers DATE, EVENT, PEOPLE, COMMENT, INFORMATION SOURCE has been initiated and currently has over 90 entries. While this is not definitive historical research, it is a quick look at the evolution of Genesis curation with pointers to documents or people with information. Engineers for future missions may find useful points of comparison for development of facilities. References:[1] Genesis Mission Reference Document, (2011) JPL D-62382.[2] Niebur S. M., edited by Brown D. W. (2023) NASA’s Discovery Program: The First 20 Years of Competitive Planetary Exploration, NASA-SP-2023-4238.[3] Genesis Mishap Investigation Board Report, Vol. 1 (July 2005).

solar wind↗

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian↗

Aerospace Power Systems Design and Analysis (APSDA) Tool

The conceptual design of space and/or planetary electrical power systems has required considerable effort. Traditionally, in the early stages of the design cycle (conceptual design), the researchers have had to thoroughly study and analyze tradeoffs between system components, hardware architectures, and operating parameters (such as frequencies) to optimize system mass, efficiency, reliability, and cost. This process could take anywhere from several months to several years (as for the former Space Station Freedom), depending on the scale of the system. Although there are many sophisticated commercial software design tools for personal computers (PC's), none of them can support or provide total system design. To meet this need, researchers at the NASA Lewis Research Center cooperated with Professor George Kusic from the University of Pittsburgh to develop a new tool to help project managers and design engineers choose the best system parameters as quickly as possible in the early design stages (in days instead of months). It is called the Aerospace Power Systems Design and Analysis (APSDA) Tool. By using this tool, users can obtain desirable system design and operating parameters such as system weight, electrical distribution efficiency, bus power, and electrical load schedule. With APSDA, a large-scale specific power system was designed in a matter of days. It is an excellent tool to help designers make tradeoffs between system components, hardware architectures, and operation parameters in the early stages of the design cycle. user interface. It operates on any PC running the MS-DOS (Microsoft Corp.) operating system, version 5.0 or later. A color monitor (EGA or VGA) and two-button mouse are required. The APSDA tool was presented at the 30th Intersociety Energy Conversion Engineering Conference (IECEC) and is being beta tested at several NASA centers. Beta test packages are available for evaluation by contacting the author.

Truong, Long V.↗

A Proposed Strategy for the U.S. to Develop and Maintain a Mainstream Capability Suite ("Warehouse") for Automated/Autonomous Rendezvous and Docking in Low Earth Orbit and Beyond

The ability of space assets to rendezvous and dock/capture/berth is a fundamental enabler for numerous classes of NASA fs missions, and is therefore an essential capability for the future of NASA. Mission classes include: ISS crew rotation, crewed exploration beyond low-Earth-orbit (LEO), on-orbit assembly, ISS cargo supply, crewed satellite servicing, robotic satellite servicing / debris mitigation, robotic sample return, and robotic small body (e.g. near-Earth object, NEO) proximity operations. For a variety of reasons to be described, NASA programs requiring Automated/Autonomous Rendezvous and Docking/Capture/Berthing (AR&D) capabilities are currently spending an order-of-magnitude more than necessary and taking twice as long as necessary to achieve their AR&D capability, "reinventing the wheel" for each program, and have fallen behind all of our foreign counterparts in AR&D technology (especially autonomy) in the process. To ensure future missions' reliability and crew safety (when applicable), to achieve the noted cost and schedule savings by eliminate costs of continually "reinventing the wheel ", the NASA AR&D Community of Practice (CoP) recommends NASA develop an AR&D Warehouse, detailed herein, which does not exist today. The term "warehouse" is used herein to refer to a toolbox or capability suite that has pre-integrated selectable supply-chain hardware and reusable software components that are considered ready-to-fly, low-risk, reliable, versatile, scalable, cost-effective, architecture and destination independent, that can be confidently utilized operationally on human spaceflight and robotic vehicles over a variety of mission classes and design reference missions, especially beyond LEO. The CoP also believes that it is imperative that NASA coordinate and integrate all current and proposed technology development activities into a cohesive cross-Agency strategy to produce and utilize this AR&D warehouse. An initial estimate indicates that if NASA strategically coordinates the development of a robust AR&D capability across the Agency, the cost of implementing AR&D on a spacecraft could be reduced from roughly $70M per mission to as low as $7M per mission, and the associated development time could be reduced from 4 years to 2 years, after the warehouse is completely developed. Table 1 shows the clear long-term benefits to the Agency in term of costs and schedules for various missions. (The methods used to arrive at the Table 1 numbers is presented in Appendices A and B.)

Krishnakumar, Kalmanje S.↗

Effect of system workload on operating system reliability - A study on IBM 3081

This paper presents an analysis of operating system failures on an IBM 3081 running VM/SP. Three broad categories of software failures are found: error handling, program control or logic, and hardware related; it is found that more than 25 percent of software failures occur in the hardware/software interface. Measurements show that results on software reliability cannot be considered representative unless the system workload is taken into account. The overall CPU execution rate, although measured to be close to 100 percent most of the time, is not found to correlate strongly with the occurrence of failures. Possible reasons for the observed workload failure dependency, based on detailed investigations of the failure data, are discussed.

Iyer, R. K.↗

DEPEND - A design environment for prediction and evaluation of system dependability

The development of DEPEND, an integrated simulation environment for the design and dependability analysis of fault-tolerant systems, is described. DEPEND models both hardware and software components at a functional level, and allows automatic failure injection to assess system performance and reliability. It relieves the user of the work needed to inject failures, maintain statistics, and output reports. The automatic failure injection scheme is geared toward evaluating a system under high stress (workload) conditions. The failures that are injected can affect both hardware and software components. To illustrate the capability of the simulator, a distributed system which employs a prediction-based, dynamic load-balancing heuristic is evaluated. Experiments were conducted to determine the impact of failures on system performance and to identify the failures to which the system is especially susceptible.

Goswami, Kumar K.↗

Developing Sustainable Spacecraft Water Management Systems

It is well recognized that water handling systems used in a spacecraft are prone to failure caused by biofouling and mineral scaling, which can clog mechanical systems and degrade the performance of capillary-based technologies. Long duration spaceflight applications, such as extended stays at a Lunar Outpost or during a Mars transit mission, will increasingly benefit from hardware that is generally more robust and operationally sustainable overtime. This paper presents potential design and testing considerations for improving the reliability of water handling technologies for exploration spacecraft. Our application of interest is to devise a spacecraft wastewater management system wherein fouling can be accommodated by design attributes of the management hardware, rather than implementing some means of preventing its occurrence.

Thomas, Evan A.↗

Flexible AI Models for Grid Resilience

The rapid growth in size and complexity of artificial intelligence (AI) and machine learning (ML) models has led to increased energy demands, posing a threat to the reliability of the existing power grid. This project addresses the challenge of highly intermittent and energy-intensive inference workloads by (1) developing fidelity-adaptive neural networks capable of dynamic response to grid conditions and (2) integrating these networks with power flow simulations to assess their impact on power grid reliability. We will explore both top-down and bottom-up approaches to create hierarchies of submodels that provide a controlled trade-off between power draw and prediction accuracy. The top-down method utilizes NN pruning to reduce a flagship model into progressively smaller, energy-efficient variants. The bottom-up approach employs geometrically principled weight setting strategies to construct depth-efficient models from the ground up. A real-time hardware-in-the-loop (HIL) platform will be developed to simulate a scaled AC power grid, integrating live AI workload power draw and enabling dynamic model switching in response to grid feedback. This work will provide a novel framework for evaluating the impact of flexible AI/ML workloads on grid performance and establish new methodologies for energy-aware computing in data centers. The outcomes will demonstrate that adaptive AI/ML can play a critical role in improving grid stability while advancing NREL's leadership in energy-efficient computing research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Thermal cycling testing to failure of a ceramic column grid array package for space applications

Ceramic column grid array (CCGA) packages have been used increasingly in logic and microprocessor functions, telecommunications, flight avionics boards, payload electronics boards, engineering navigational and science cameras, electronic assemblies, and other applications, because of their inherent advantages such as high electrical interconnect density, very good thermal and electrical performance, and compatibility with standard surface-mount technology (SMT) packaging assembly processes. Because these advanced electronic packages tend to have less solder-joint strain-relief than do leaded flat-pack electronic packages, the reliability of CCGA packages in challenging thermal environments is a very important consideration for their short- and long-term use in JPL-NASA space missions. In this study, we assembled daisy chains of polyimide printed wiring boards from CCGA-interconnect packages, inspected the boards nondestructively, and then subjected them to thermal cycling to assess their reliability in thermal environments from +125°C to -40°C±25°C. The test hardware consists of a CCGA 1752 package (CN version). The package was divided into four daisy-chained sections that were electrically monitored for their continuity during thermal cycling. The CCGA 1752 package is roughly 45 mm × 45 mm with a 42-mm × 42-mm array of 80%/20% Pb/Sn columns on a 1.00-mm pitch. The resistance of the daisy-chained CCGA interconnects was continuously monitored during thermal cycling in a gaseous nitrogen environment. Electrical continuity resistance measurements as a function of thermal cycling are reported here; tests to date have shown significant change to an open circuit in daisy-chain resistance as a function of thermal cycling. The change in interconnect resistance becomes increasingly noticeable with increasing number of thermal cycles. This paper describes the experimental thermal-cycling test results of CCGA 1752 package reliability testing under an extremely wide temperature range. The first failure was observed at 1479th thermal cycle. We report the thermal-cycle reliability test data for ~2500 thermal cycles.

Ramesham, Rajeshuni↗

Sublimator Driven Coldplate Engineering Development Unit Test Results and Development of Second Generation SDC

The Sublimator Driven Coldplate (SDC) is a unique piece of thermal control hardware that has several advantages over a traditional thermal control scheme. The principal advantage is the possible elimination of a pumped fluid loop, potentially increasing reliability and reducing complexity while saving both mass and power. Furthermore, the Integrated Sublimator Driven Coldplate (ISDC) concept couples a coolant loop with the previously described SDC hardware. This combination allows the SDC to be used as a traditional coldplate during long mission phases. The previously developed SDC technology cannot be used for long mission phases due to the fact that it requires a consumable feedwater for heat rejection. Adding a coolant loop also provides for dissimilar redundancy on the Altair Lander ascent module thermal control system, which is the target application for this technology. Tests were performed on an Engineering Development Unit at NASA s Johnson Space Center to quantify and assess the performance of the SDC. Correlated thermal math models were developed to help explain the test data. The paper also outlines the preliminary results of an ISDC concept being developed.

Stephan, Ryan A.↗