Search NASA⌕ Search

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34

A Graph Neural Network Surrogate Model for hls4ml

Recent advancements in use of machine learning (ML) techniques on field-programmable gate arrays (FPGAs) have allowed for the implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area are strictly bounded. The hls4ml framework is a procedure that converts trained ML model software to a synthesis result to can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it may not be possible to successfully convert a model into a synthesis result, or the resource consumption of the model may exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model using a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of a given model when passed through the hls4ml pipeline, without needing to run the pipeline.

Plotnikov, Dennis↗

Optimization Testbed Cometboards Extended into Stochastic Domain

COMparative Evaluation Testbed of Optimization and Analysis Routines for the Design of Structures (CometBoards) is a multidisciplinary design optimization software. It was originally developed for deterministic calculation. It has now been extended into the stochastic domain for structural design problems. For deterministic problems, CometBoards is introduced through its subproblem solution strategy as well as the approximation concept in optimization. In the stochastic domain, a design is formulated as a function of the risk or reliability. Optimum solution including the weight of a structure, is also obtained as a function of reliability. Weight versus reliability traced out an inverted-S-shaped graph. The center of the graph corresponded to 50 percent probability of success, or one failure in two samples. A heavy design with weight approaching infinity could be produced for a near-zero rate of failure that corresponded to unity for reliability. Weight can be reduced to a small value for the most failure-prone design with a compromised reliability approaching zero. The stochastic design optimization (SDO) capability for an industrial problem was obtained by combining three codes: MSC/Nastran code was the deterministic analysis tool, fast probabilistic integrator, or the FPI module of the NESSUS software, was the probabilistic calculator, and CometBoards became the optimizer. The SDO capability requires a finite element structural model, a material model, a load model, and a design model. The stochastic optimization concept is illustrated considering an academic example and a real-life airframe component made of metallic and composite materials.

Patnaik, Surya N.↗

Mission Operations Center (MOC) - Precipitation Processing System (PPS) Interface Software System (MPISS)

MPISS is an automatic file transfer system that implements a combination of standard and mission-unique transfer protocols required by the Global Precipitation Measurement Mission (GPM) Precipitation Processing System (PPS) to control the flow of data between the MOC and the PPS. The primary features of MPISS are file transfers (both with and without PPS specific protocols), logging of file transfer and system events to local files and a standard messaging bus, short term storage of data files to facilitate retransmissions, and generation of file transfer accounting reports. The system includes a graphical user interface (GUI) to control the system, allow manual operations, and to display events in real time. The PPS specific protocols are an enhanced version of those that were developed for the Tropical Rainfall Measuring Mission (TRMM). All file transfers between the MOC and the PPS use the SSH File Transfer Protocol (SFTP). For reports and data files generated within the MOC, no additional protocols are used when transferring files to the PPS. For observatory data files, an additional handshaking protocol of data notices and data receipts is used. MPISS generates and sends to the PPS data notices containing data start and stop times along with a checksum for the file for each observatory data file transmitted. MPISS retrieves the PPS generated data receipts that indicate the success or failure of the PPS to ingest the data file and/or notice. MPISS retransmits the appropriate files as indicated in the receipt when required. MPISS also automatically retrieves files from the PPS. The unique feature of this software is the use of both standard and PPS specific protocols in parallel. The advantage of this capability is that it supports users that require the PPS protocol as well as those that do not require it. The system is highly configurable to accommodate the needs of future users.

Ferrara, Jeffrey↗

Results of the Systems Autonomy Demonstration Project

The Systems Autonomy Demonstration Project (SADP) produced a knowledge-based real-time control system for real-time control and fault detection, isolation and recovery (FDIR) of a prototype two-phase Space Station Freedom external control system (TCS). The Thermal Expert System (TEXSYS) was demonstrated in recent tests to be capable of reliable fault anticipation and detection, as well as nominal control of the thermal bus. Performance requirements were addressed by adopting a hierarchical symbolic control approach, layering model-based expert system software on a conventional numerical data acquisition and control system. The model-based capabilities of TEXSYS were shown to be advantageous, particularly for detection of unforeseen faults and sensor failures.

Glass, B. J.↗

Neural networks: Alternatives to conventional techniques for automatic docking

Automatic docking of orbiting spacecraft is a crucial operation involving the identification of vehicle orientation as well as complex approach dynamics. The chaser spacecraft must be able to recognize the target spacecraft within a scene and achieve accurate closing maneuvers. In a video-based system, a target scene must be captured and transformed into a pattern of pixels. Successful recognition lies in the interpretation of this pattern. Due to their powerful pattern recognition capabilities, artificial neural networks offer a potential role in interpretation and automatic docking processes. Neural networks can reduce the computational time required by existing image processing and control software. In addition, neural networks are capable of recognizing and adapting to changes in their dynamic environment, enabling enhanced performance, redundancy, and fault tolerance. Most neural networks are robust to failure, capable of continued operation with a slight degradation in performance after minor failures. This paper discusses the particular automatic docking tasks neural networks can perform as viable alternatives to conventional techniques.

Vinz, Bradley L.↗

Risk as a Resource - A New Paradigm

NASA must change dramatically because of the current United States federal budget climate. The American people and their elected officials have mandated a smaller, more efficient and effective government. For the past decade, NASA's budget had grown at or slightly above the rate of inflation. In that era, taking all steps to avoid the risk of failure was the rule. Spacecraft development was characterized by extensive analyses, numerous reviews, and multiple conservative tests. This methodology was consistent with the long available schedules for developing hardware and software for very large, billion dollar spacecraft. Those days are over. The time when every identifiable step was taken to avoid risk is being replaced by a new paradigm which manages risk in much the same way as other resources (schedule, performance, or dollars) are managed. While success is paramount to survival, it can no longer be bought with a large growing NASA budget.

Risk Resource↗

Verification and Validation of Adaptive and Intelligent Systems with Flight Test Results

F-15 IFCS project goals are: a) Demonstrate Control Approaches that can Efficiently Optimize Aircraft Performance in both Normal and Failure Conditions [A] & [B] failures. b) Advance Neural Network-Based Flight Control Technology for New Aerospace Systems Designs with a Pilot in the Loop. Gen II objectives include; a) Implement and Fly a Direct Adaptive Neural Network Based Flight Controller; b) Demonstrate the Ability of the System to Adapt to Simulated System Failures: 1) Suppress Transients Associated with Failure; 2) Re-Establish Sufficient Control and Handling of Vehicle for Safe Recovery. c) Provide Flight Experience for Development of Verification and Validation Processes for Flight Critical Neural Network Software.

Burken, John J.↗

Orion MPCV GN and C End-to-End Phasing Tests

End-to-end integration tests are critical risk reduction efforts for any complex vehicle. Phasing tests are an end-to-end integrated test that validates system directional phasing (polarity) from sensor measurement through software algorithms to end effector response. Phasing tests are typically performed on a fully integrated and assembled flight vehicle where sensors are stimulated by moving the vehicle and the effectors are observed for proper polarity. Orion Multi-Purpose Crew Vehicle (MPCV) Pad Abort 1 (PA-1) Phasing Test was conducted from inertial measurement to Launch Abort System (LAS). Orion Exploration Flight Test 1 (EFT-1) has two end-to-end phasing tests planned. The first test from inertial measurement to Crew Module (CM) reaction control system thrusters uses navigation and flight control system software algorithms to process commands. The second test from inertial measurement to CM S-Band Phased Array Antenna (PAA) uses navigation and communication system software algorithms to process commands. Future Orion flights include Ascent Abort Flight Test 2 (AA-2) and Exploration Mission 1 (EM-1). These flights will include additional or updated sensors, software algorithms and effectors. This paper will explore the implementation of end-to-end phasing tests on a flight vehicle which has many constraints, trade-offs and compromises. Orion PA-1 Phasing Test was conducted at White Sands Missile Range (WSMR) from March 4-6, 2010. This test decreased the risk of mission failure by demonstrating proper flight control system polarity. Demonstration was achieved by stimulating the primary navigation sensor, processing sensor data to commands and viewing propulsion response. PA-1 primary navigation sensor was a Space Integrated Inertial Navigation System (INS) and Global Positioning System (GPS) (SIGI) which has onboard processing, INS (3 accelerometers and 3 rate gyros) and no GPS receiver. SIGI data was processed by GN&C software into thrust magnitude and direction commands. The processing changes through three phases of powered flight: pitchover, downrange and reorientation. The primary inputs to GN&C are attitude position, attitude rates, angle of attack (AOA) and angle of sideslip (AOS). Pitch and yaw attitude and attitude rate responses were verified by using a flight spare SIGI mounted to a 2-axis rate table. AOA and AOS responses were verified by using a data recorded from SIGI movements on a robotic arm located at NASA Johnson Space Center. The data was consolidated and used in an open-loop data input to the SIGI. Propulsion was the Launch Abort System (LAS) Attitude Control Motor (ACM) which consisted of a solid motor with 8 nozzles. Each nozzle has active thrust control by varying throat area with a pintle. LAS ACM pintles are observable through optically transparent nozzle covers. SIGI movements on robot arm, SIGI rate table movements and LAS ACM pintle responses were video recorded as test artifacts for analysis and evaluation. The PA-1 Phasing Test design was determined based on test performance requirements, operational restrictions and EGSE capabilities. This development progressed during different stages. For convenience these development stages are initial, working group, tiger team, Engineering Review Team (ERT) and final.

Neumann, Brian C.↗

Independent Orbiter Assessment (IOA): Analysis of the backup flight system

The results of the Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis (FMEA) and Critical Items List (CIL) are presented. The IOA approach features a top-down analysis of the hardware to determine failure modes, criticality, and potential critical items. To preserve independence, this analysis was accomplished without reliance upon the results contained within the NASA FMEA/CIL documentation. This report documents the analysis results corresponding to the Orbiter Backup Flight System (BFS) hardware. The BFS hardware consists of one General Purpose Computer (GPC) loaded with backup flight software and the components used to engage/disengage that unique GPC. Specifically, the BFS hardware includes the following: DDU (Display Driver Unit), BFC (Backup Flight Controller), GPC (General Purpose Computer), switches (engage, disengage, GPC, CRT), and circuit protectors (fuses, circuit breakers). The IOA analysis process utilized available BFS hardware drawings and schematics for defining hardware assemblies, components, and hardware items. Each level of hardware was evaluated and analyzed for possible failure modes and effects. Criticality was assigned based upon the severity of the effect for each failure mode. Of the failure modes analyzed, 19 could potentially result in a loss of life and/or loss of vehicle.

Prust, E. E.↗

Fault Management Techniques in Human Spaceflight Operations

This paper discusses human spaceflight fault management operations. Fault detection and response capabilities available in current US human spaceflight programs Space Shuttle and International Space Station are described while emphasizing system design impacts on operational techniques and constraints. Preflight and inflight processes along with products used to anticipate, mitigate and respond to failures are introduced. Examples of operational products used to support failure responses are presented. Possible improvements in the state of the art, as well as prioritization and success criteria for their implementation are proposed. This paper describes how the architecture of a command and control system impacts operations in areas such as the required fault response times, automated vs. manual fault responses, use of workarounds, etc. The architecture includes the use of redundancy at the system and software function level, software capabilities, use of intelligent or autonomous systems, number and severity of software defects, etc. This in turn drives which Caution and Warning (C&W) events should be annunciated, C&W event classification, operator display designs, crew training, flight control team training, and procedure development. Other factors impacting operations are the complexity of a system, skills needed to understand and operate a system, and the use of commonality vs. optimized solutions for software and responses. Fault detection, annunciation, safing responses, and recovery capabilities are explored using real examples to uncover underlying philosophies and constraints. These factors directly impact operations in that the crew and flight control team need to understand what happened, why it happened, what the system is doing, and what, if any, corrective actions they need to perform. If a fault results in multiple C&W events, or if several faults occur simultaneously, the root cause(s) of the fault(s), as well as their vehicle-wide impacts, must be determined in order to maintain situational awareness. This allows both automated and manual recovery operations to focus on the real cause of the fault(s). An appropriate balance must be struck between correcting the root cause failure and addressing the impacts of that fault on other vehicle components. Lastly, this paper presents a strategy for using lessons learned to improve the software, displays, and procedures in addition to determining what is a candidate for automation. Enabling technologies and techniques are identified to promote system evolution from one that requires manual fault responses to one that uses automation and autonomy where they are most effective. These considerations include the value in correcting software defects in a timely manner, automation of repetitive tasks, making time critical responses autonomous, etc. The paper recommends the appropriate use of intelligent systems to determine the root causes of faults and correctly identify separate unrelated faults.

O'Hagan, Brian↗

Pilot interaction with automated airborne decision making systems

Progress was made in the three following areas. In the rule-based modeling area, two papers related to identification and significane testing of rule-based models were presented. In the area of operator aiding, research focused on aiding operators in novel failure situations; a discrete control modeling approach to aiding PLANT operators was developed; and a set of guidelines were developed for implementing automation. In the area of flight simulator hardware and software, the hardware will be completed within two months and initial simulation software will then be integrated and tested.

Rouse, W. B.↗

Advanced Information Processing System (AIPS)

Advanced Information Processing System (AIPS) is a computer systems philosophy, a set of validated hardware building blocks, and a set of validated services as embodied in system software. The goal of AIPS is to provide the knowledgebase which will allow achievement of validated fault-tolerant distributed computer system architectures, suitable for a broad range of applications, having failure probability requirements of 10E-9 at 10 hours. A background and description is given followed by program accomplishments, the current focus, applications, technology transfer, FY92 accomplishments, and funding.

Pitts, Felix L.↗

RICIS Symposium 1992: Mission and Safety Critical Systems Research and Applications

This conference deals with computer systems which control systems whose failure to operate correctly could produce the loss of life and or property, mission and safety critical systems. Topics covered are: the work of standards groups, computer systems design and architecture, software reliability, process control systems, knowledge based expert systems, and computer and telecommunication protocols.

Source record↗

CARES/Life Software for Designing More Reliable Ceramic Parts

Products made from advanced ceramics show great promise for revolutionizing aerospace and terrestrial propulsion, and power generation. However, ceramic components are difficult to design because brittle materials in general have widely varying strength values. The CAPES/Life software eases this task by providing a tool to optimize the design and manufacture of brittle material components using probabilistic reliability analysis techniques. Probabilistic component design involves predicting the probability of failure for a thermomechanically loaded component from specimen rupture data. Typically, these experiments are performed using many simple geometry flexural or tensile test specimens. A static, dynamic, or cyclic load is applied to each specimen until fracture. Statistical strength and SCG (fatigue) parameters are then determined from these data. Using these parameters and the results obtained from a finite element analysis, the time-dependent reliability for a complex component geometry and loading is then predicted. Appropriate design changes are made until an acceptable probability of failure has been reached.

Nemeth, Noel N.↗

Making or Breaking a Rover: System Engineering Parameters On-Board the Mars 2020 Perseverance Rover

On February 18, 2021, Perseverance, NASA’s Jet Propulsion Laboratory’s (JPL’s) Mars 2020 Rover, successfully landed on Mars with all systems nominal, despite the risk surrounding the over 200,000 internal flight parameters that had to be properly configured. The Perseverance team defines these parameters as software variables that are configurable, commandable and retrievable from Earth. In 2015, the Mars 2020 project leaders focused on improving systems engineering of parameters based on their experiences from parameter management on previous Mars rovers (Curiosity, Opportunity, Spirit, and Pathfinder) and parameter failures of past missions, such as the mission-ending parameter of the Mars Climate Orbiter. The new rigorous development process allowed for efficient certification and effective implementation of the parameters, allowing the rover to approach and land on the red planet (the most challenging phase of the mission) with zero parameter issues. Although successful, the Perseverance team learned many lessons for how to better manage parameters for the continued surface operations of the Mars 2020 mission and future missions. This paper will discuss eight parameter-management topics for the Perseverance Mission. The first is parameter definition: how we define parameters on our mission, where they are physically located on the vehicle, and why we have so many of them. The second topic is the updated parameter flight software module from Curiosity, including details on the 99% reduction in parameter commands, new bulk configuration capabilities, and improved parameter traceability. The third topic is parameter selection for different mission phases; this includes improving and tweaking our preferred parameter settings until they become certification candidates and managing parameter configurations based on test venue throughout the mission life cycle. The fourth topic is our flight certification process; this includes certification of flight values for four different epochs in the mission: Launch, Entry Decent and Landing (EDL) - 6days, Landing + 5 Sols (Martian Days, still on Cruise Flight Software), and once are on Surface Flight Software (FSW). The fifth topic covers in-flight command implementation, along with details on testing, validation, and verification of those commands. In the sixth section, we will explain our use of open-source management tools, including how we used GitHub for version control and management approvals. The seventh topic will describe the ground tools used in operations, including capabilities of the in-house built tool called Parasol. The eighth and final topic will dig into lessons learned for improving parameter management in the future of this mission and others.

Roth, Brian↗

Making or Breaking a Rover- Systems Engineering Parameters On-Board the Mars 2020 Perseverance Rover

On February 18, 2021, Perseverance, NASA’s Jet Propulsion Laboratory’s (JPL’s) Mars 2020 Rover, successfully landed on Mars with all systems nominal, despite the risk surrounding the over 200,000 internal flight parameters that had to be properly configured. The Perseverance team defines these parameters as software variables that are configurable, commandable and retrievable from Earth. In 2015, the Mars 2020 project leaders focused on improving systems engineering of parameters based on their experiences from parameter management on previous Mars rovers (Curiosity, Opportunity, Spirit, and Pathfinder) and parameter failures of past missions, such as the mission-ending parameter of the Mars Climate Orbiter. The new rigorous development process allowed for efficient certification and effective implementation of the parameters, allowing the rover to approach and land on the red planet (the most challenging phase of the mission) with zero parameter issues. Although successful, the Perseverance team learned many lessons for how to better manage parameters for the continued surface operations of the Mars 2020 mission and future missions. This paper will discuss eight parameter-management topics for the Perseverance Mission. The first is parameter definition: how we define parameters on our mission, where they are physically located on the vehicle, and why we have so many of them. The second topic is the updated parameter flight software module from Curiosity, including details on the 99% reduction in parameter commands, new bulk configuration capabilities, and improved parameter traceability. The third topic is parameter selection for different mission phases; this includes improving and tweaking our preferred parameter settings until they become certification candidates and managing parameter configurations based on test venue throughout the mission life cycle. The fourth topic is our flight certification process; this includes certification of flight values for four different epochs in the mission: Launch, Entry Decent and Landing (EDL) - 6days, Landing + 5 Sols (Martian Days, still on Cruise Flight Software), and once are on Surface Flight Software (FSW). The fifth topic covers in-flight command implementation, along with details on testing, validation, and verification of those commands. In the sixth section, we will explain our use of open-source management tools, including how we used GitHub for version control and management approvals. The seventh topic will describe the ground tools used in operations, including capabilities of the in-house built tool called Parasol. The eighth and final topic will dig into lessons learned for improving parameter management in the future of this mission and others.

Roth, Brian↗

Seismic Contingency Auto Generator

This code takes in premade earthquake scenario XML files from USGS, power grid data, and converts them into a contingency file (.con file) that can be used by power grid solvers. Within the .con file are a number (Specified by the user) of contingencies that have randomly failed power transformers based on their likelihood of failure and peak ground acceleration (PGA) value around the transformer. The transformers' likelihood of failure was calculated based on a variety of finite element modeling on various transformer designed for specific transformer voltage classes. Parameters from these FEM were used to create generic fragility curves for transformers within a specific voltage class, which correspond with earthquake PGA values to produced a probability of failure for a given earthquake scenario. More refined versions of this process, such as specifying specific transformer design categories within a voltage class, could also be applied in future iterations of the software.

Vaagensmith, Bjorn [Idaho National Laboratory (INL↗

Develop advanced nonlinear signal analysis topographical mapping system

This study will provide timely assessment of SSME component operational status, identify probable causes of malfunction, and indicate feasible engineering solutions. The final result of this program will yield an advanced nonlinear signal analysis topographical mapping system (ATMS) of nonlinear and nonstationary spectral analysis software package integrated with the Compressed SSME TOPO Data Base (CSTDB) on the same platform. This system will allow NASA engineers to retrieve any unique defect signatures and trends associated with different failure modes and anomalous phenomena over the entire SSME test history across turbopump families.

Jong, Jen-Yi↗