Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software Reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

ICE-RASSOR: Intelligent Capabilities Enhanced Regolith Advanced Surface Systems Operations Robot

NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU) processing. RASSOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar surface, RASSOR software and sensory systems need to be robust and maximize the information extracted from on-board sensory. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We apply supervised learning using real data to estimate the soil mass collected without the need for mass flow rate monitors or other explicate sensing techniques. We also create a reduced-order simulation environment to develop autonomous trenching controllers via reinforcement learning and proto-type state estimation architectures. Our initial results suggest that excavated regolith mass can be inferred within 2.9% RMS error of full scale, and reinforcement learning for autonomous operations has learned viable trenching strategies and helped identify desirable sensing capabilities, arrangements, and considerations. Future work includes regolith mass estimation during dynamic operation, expanding our simulation to more complex environments, and transfer learning from simulation to hardware.

machine learning↗

Streamlining GNC Architecture Development and FSW Integration for the Mars Ascent Vehicle

The Mars Ascent Vehicle (MAV) will be the first vehicle to perform an ascent from the surface of another atmospheric planetary body outside of the Earth-Moon system. Significant light-time delay requires complete autonomy of flight throughout ascent, and naturally a high level of reliability is desired in both MAV’s hardware and software subsystems. The MAV Guidance, Navigation and Controls (GNC) team and the MAV Flight Software (FSW) team have partnered together to improve the efficiency of algorithm integration onto the MAV flight processor, and to increase confidence that said integration is successful and without human error. An interface architecture is proposed for the GNC suite that allows both the guidance and navigation subsystems to provide code algorithms directly in C++, and the controls subsystem to provide MATLAB Simulink auto-coded algorithms. Several continuous integration/deployment (CI/CD) methodologies have been considered for ease of transition of algorithm code from the GNC team to the FSW team. The GNC/FSW teams also worked together to develop a cFS-friendly wrapper which abstracts the integration of the GNC algorithm code into an interface-level API that is compatible with cFS. Several iterations of vehicle GNC code have been produced between the GNC/FSW team’s partnership, and this strong interface between these two teams have allowed the GNC/FSW teams to greatly increase confidence of efficient and error-free implementation of the GNC code onto MAV for a successful flight.

Engineering↗

Towards Autonomous Lunar Resource Excavation via Reinforcement Learning

To continue on a sustainable and flexible path, NASA needs to address the challenge of collecting and moving large amounts of regolith at the destination. NASA’s Regolith Advanced Surface Systems Operations Robot (RASSOR) is principally designed to mine and deliver regolith for In-Situ Resource Utilization (ISRU) processing. RASSOR’s design enables it to efficiently collect and deposit regolith, return collected material for processing, and myriad related ISRU activities. To reliably perform these operations on the lunar surface, RASSOR software and sensory systems need to be robust and maximize the information extracted from a reduced sensor payload. Herein, we present preliminary findings from the Intelligent Capabilities Enhanced RASSOR project. We created reduced-order simulation environments to develop autonomous trenching controllers via reinforcement learning and prototype state estimation architectures. The goal of reinforcement learning is for an agent to learn a policy (task strategy) through interactions with an environment. When the agent performs an action, a change occurs in environment state and a numerical reward is received which informs the agent whether the action performed was good or not. Since reinforcement learning algorithms learn through trial-and-error, a simulation is a desirable first environment for development and learning. We developed two simulations, the first is a 2D excavation simulation developed to facilitate parameter selection, and a 3D simulation developed using a game physics engine, to simulate simplified soil interactions and increase the fidelity of the dynamic models of the robotic agents. The development of this 3D simulation has enabled the training of additional sensing capabilities and research both at the granular mechanics and operations levels. We experimented with various virtual sensor payloads to identify a combination that enabled efficient excavation operation and learning. Our reward function is based on how much material is excavated per step. A penalty is also received for leaving the dig site and to smooth the acceleration of the drum arms. We implemented pseudo time-of-flight sensors to report distance from each drum to ground and the height above ground which was found to be more efficient than existing solutions. Our findings suggest that reinforcement learning for autonomous operations has learned viable trenching strategies within 3000 training episodes in our simplified 2D environment and helped identify desirable sensing capabilities, arrangements, and considerations such as the positioning of time-of-flight sensors. Future work includes expanding our simulation to more complex environments and scenarios, and transfer learning from simulation to RASSOR 2.0 hardware for deployment in the Regolith Test Bin at NASA's Kennedy Space Center.

rassor↗

Providing Experimental Infrastructure for Accelerating Advanced Reactor Demonstrations through the National Reactor Innovation Center

A suite of experimental infrastructure projects has been developed by the National Reactor Innovation Center to accelerate advanced reactor demonstrations and facilitate their development, addressing crucial gaps in data, materials characterization, and modeling. First, the Molten Salt Thermophysical Examination Capability (MSTEC) provides a specialized platform for post-irradiation characterization of molten salt reactor fuel, coolant salts, and structural materials, essential for supporting the design and operation of advanced reactors and future commercial molten salt reactor development and licensing. The Virtual Test Bed (VTB) complements these efforts by leveraging advanced modeling and simulation tools to evaluate reactor performance and safety. Serving as a library of reference models, the VTB offers a database of multiphysics reactor models, facilitating rapid safety evaluations and includes continuous software quality assurance, crucial for accelerating deployment while maintaining reliability. Additionally, the Helium Component Test Facility (HeCTF) addresses the need for high-temperature helium-cooled reactor component testing. As the first-of-its-kind facility in the United States, HeCTF emulates high-temperature gas reactor conditions, reducing time and cost associated with component validation, thereby accelerating reactor development. Finally, In-cell Thermal Creep Frames provide a unique solution for obtaining thermal creep data from irradiated materials, critical for materials qualification and licensing. Developed by the National Reactor Innovation Center, these compact frames enable the examination of previously irradiated materials, overcoming traditional limitations and enhancing the understanding of mechanical properties crucial for reactor development. Collectively, these experimental infrastructure projects form a comprehensive framework aimed at expediting advanced reactor demonstrations, fostering innovation, and ensuring the viability of next-generation nuclear energy solutions.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Validation of highly reliable, real-time knowledge-based systems

Knowledge-based systems have the potential to greatly increase the capabilities of future aircraft and spacecraft and to significantly reduce support manpower needed for the space station and other space missions. However, a credible validation methodology must be developed before knowledge-based systems can be used for life- or mission-critical applications. Experience with conventional software has shown that the use of good software engineering techniques and static analysis tools can greatly reduce the time needed for testing and simulation of a system. Since exhaustive testing is infeasible, reliability must be built into the software during the design and implementation phases. Unfortunately, many of the software engineering techniques and tools used for conventional software are of little use in the development of knowledge-based systems. Therefore, research at Langley is focused on developing a set of guidelines, methods, and prototype validation tools for building highly reliable, knowledge-based systems. The use of a comprehensive methodology for building highly reliable, knowledge-based systems should significantly decrease the time needed for testing and simulation. A proven record of delivering reliable systems at the beginning of the highly visible testing and simulation phases is crucial to the acceptance of knowledge-based systems in critical applications.

Johnson, Sally C.↗

Ultra reliability at NASA

Ultra reliable systems are critical to NASA particularly as consideration is being given to extended lunar missions and manned missions to Mars. NASA has formulated a program designed to improve the reliability of NASA systems. The long term goal for the NASA ultra reliability is to ultimately improve NASA systems by an order of magnitude. The approach outlined in this presentation involves the steps used in developing a strategic plan to achieve the long term objective of ultra reliability. Consideration is given to: complex systems, hardware (including aircraft, aerospace craft and launch vehicles), software, human interactions, long life missions, infrastructure development, and cross cutting technologies. Several NASA-wide workshops have been held, identifying issues for reliability improvement and providing mitigation strategies for these issues. In addition to representation from all of the NASA centers, experts from government (NASA and non-NASA), universities and industry participated. Highlights of a strategic plan, which is being developed using the results from these workshops, will be presented.

risk↗

Multi-Agent Control Planes for Quantum Networks: A Scalable Architecture for Autonomous Quantum Internet Management

Quantum networks are expected to enable distributed quantum computing, secure communication, and global entanglement distribution. However, operating such networks presents significant challenges, including stochastic quantum processes, fragile entanglement resources, dynamic topology, and cross-layer control requirements. Current quantum network control architectures largely rely on centralized or hierarchical controllers inspired by classical software-defined networking (SDN). While effective for small testbeds, these approaches face scalability, latency, and reliability limitations as quantum networks grow. This paper proposes a multi-agent control plane architecture for quantum networks. In this design, intelligent software agents operate at quantum nodes, repeaters, and orchestration layers, collectively managing entanglement generation, routing, purification, and scheduling. The distributed intelligence of the agent system allows the network to adapt dynamically to quantum hardware variability and environmental noise. We argue that multi-agent systems provide significant advantages over centralized control approaches, including scalability, resilience, local autonomy, and real-time adaptation. The paper discusses architectural design principles, agent coordination mechanisms, and research challenges in deploying multi-agent control planes for the emerging quantum Internet.

Alnajjar, Anees [ORNL] (ORCID:0000000237101601)↗

The software-implemented fault tolerance /SIFT/ approach to fault tolerant computing

SIFT is an experimental computer designed for highly reliable flight-control service in advanced air transports. Its development was intended to integrate and demonstrate the latest techniques in fault-tolerant computing. During its development, several new problems of some generality were uncovered and solved. The technology developed for the validation of its design is seen as being perhaps as important as the design itself. The SIFT design is described, as is the way in which the design and its validation were shaped by the requirements of its intended application. Attention is also given to reliability and fault tolerance. The most significant feature of the hardware design is the absence of elements that can generate multiple faults, such as shared clocks or data buses. It is noted that the software is realized in only 800 lines of code, of which 80% are in a high-level language.

Goldberg, J.↗

Sensory redundancy management: The development of a design methodology for determining threshold values through a statistical analysis of sensor output data

Sensor redundancy management (SRM) requires a system which will detect failures and reconstruct avionics accordingly. A probability density function to determine false alarm rates, using an algorithmic approach was generated. Microcomputer software was developed which will print out tables of values for the cummulative probability of being in the domain of failure; system reliability; and false alarm probability, given a signal is in the domain of failure. The microcomputer software was applied to the sensor output data for various AFT1 F-16 flights and sensor parameters. Practical recommendations for further research were made.

Scalzo, F.↗

Formal Safety Certification of Aerospace Software

In principle, formal methods offer many advantages for aerospace software development: they can help to achieve ultra-high reliability, and they can be used to provide evidence of the reliability claims which can then be subjected to external scrutiny. However, despite years of research and many advances in the underlying formalisms of specification, semantics, and logic, formal methods are not much used in practice. In our opinion this is related to three major shortcomings. First, the application of formal methods is still expensive because they are labor- and knowledge-intensive. Second, they are difficult to scale up to complex systems because they are based on deep mathematical insights about the behavior of the systems (t.e., they rely on the "heroic proof"). Third, the proofs can be difficult to interpret, and typically stand in isolation from the original code. In this paper, we describe a tool for formally demonstrating safety-relevant aspects of aerospace software, which largely circumvents these problems. We focus on safely properties because it has been observed that safety violations such as out-of-bounds memory accesses or use of uninitialized variables constitute the majority of the errors found in the aerospace domain. In our approach, safety means that the program will not violate a set of rules that can range for the simple memory access rules to high-level flight rules. These different safety properties are formalized as different safety policies in Hoare logic, which are then used by a verification condition generator along with the code and logical annotations in order to derive formal safety conditions; these are then proven using an automated theorem prover. Our certification system is currently integrated into a model-based code generation toolset that generates the annotations together with the code. However, this automated formal certification technology is not exclusively constrained to our code generator and could, in principle, also be integrated with other code generators such as RealTime Workshop or even applied to legacy code. Our approach circumvents the historical problems with formal methods by increasing the degree of automation on all levels. The restriction to safety policies (as opposed to arbitrary functional behavior) results in simpler proof problems that can generally be solved by fully automatic theorem proves. An automated linking mechanism between the safety conditions and the code provides some of the traceability mandated by process standards such as DO-178B. An automated explanation mechanism uses semantic markup added by the verification condition generator to produce natural-language explanations of the safety conditions and thus supports their interpretation in relation to the code. It shows an automatically generated certification browser that lets users inspect the (generated) code along with the safety conditions (including textual explanations), and uses hyperlinks to automate tracing between the two levels. Here, the explanations reflect the logical structure of the safety obligation but the mechanism can in principle be customized using different sets of domain concepts. The interface also provides some limited control over the certification process itself. Our long-term goal is a seamless integration of certification, code generation, and manual coding that results in a "certified pipeline" in which specifications are automatically transformed into executable code, together with the supporting artifacts necessary for achieving and demonstrating the high level of assurance needed in the aerospace domain.

Denney, Ewen↗

Optimizing Probability of Detection Point Estimate Demonstration

Probability of detection (POD) analysis is used in assessing reliably detectable flaw size in nondestructive evaluation (NDE). MIL-HDBK-18231and associated mh18232POD software gives most common methods of POD analysis. Real flaws such as cracks and crack-like flaws are desired to be detected using these NDE methods. A reliably detectable crack size is required for safe life analysis of fracture critical parts. The paper provides discussion on optimizing probability of detection (POD) demonstration experiments using Point Estimate Method. POD Point estimate method is used by NASA for qualifying special NDE procedures. The point estimate method uses binomial distribution for probability density. Normally, a set of 29 flaws of same size within some tolerance are used in the demonstration. The optimization is performed to provide acceptable value for probability of passing demonstration (PPD) and achieving acceptable value for probability of false (POF) calls while keeping the flaw sizes in the set as small as possible.

Koshti, Ajay M.↗

Software safety

Software safety and its relationship to other qualities are discussed. It is shown that standard reliability and fault tolerance techniques will not solve the safety problem for the present. A new attitude requires: looking at what you do NOT want software to do along with what you want it to do; and assuming things will go wrong. New procedures and changes to entire software development process are necessary: special software safety analysis techniques are needed; and design techniques, especially eliminating complexity, can be very helpful.

Leveson, Nancy↗

Improving and Expanding NASA Software Cost Estimation Methods

Estimators and analysts are increasingly being tasked to develop better models and reliable cost estimates in support of program planning and execution. While there has been extensive work on improving parametric methods for cost estimation, there is very little focus on the use of cost models based on analogy and clustering algorithms. In this paper we summarize the results of our research in developing an analogy method for estimating NASA spacecraft flight software using spectral clustering on system characteristics (symbolic nonnumerical data) and evaluate its performance by comparing it to a number of the most commonly used estimation methods. The strengths and weaknesses of each method based on their performance are also discussed. The paper concludes with an overview of the analogy estimation tool (ASCoT) developed for use within NASA that implements the recommended analogy algorithm.

Hihn, Jairus↗

Evolving impact of Ada on a production software environment

Many aspects of software development with Ada have evolved as our Ada development environment has matured and personnel have become more experienced in the use of Ada. The Software Engineering Laboratory (SEL) has seen differences in the areas of cost, reliability, reuse, size, and use of Ada features. A first Ada project can be expected to cost about 30 percent more than an equivalent FORTRAN project. However, the SEL has observed significant improvements over time as a development environment progresses to second and third uses of Ada. The reliability of Ada projects is initially similar to what is expected in a mature FORTRAN environment. However, with time, one can expect to gain improvements as experience with the language increases. Reuse is one of the most promising aspects of Ada. The proportion of reusable Ada software on our Ada projects exceeds the proportion of reusable FORTRAN software on our FORTRAN projects. This result was noted fairly early in our Ada projects, and experience shows an increasing trend over time.

Mcgarry, F.↗

Software development for safety-critical medical applications

There are many computer-based medical applications in which safety and not reliability is the overriding concern. Reduced, altered, or no functionality of such systems is acceptable as long as no harm is done. A precise, formal definition of what software safety means is essential, however, before any attempt can be made to achieve it. Without this definition, it is not possible to determine whether a specific software entity is safe. A set of definitions pertaining to software safety will be presented and a case study involving an experimental medical device will be described. Some new techniques aimed at improving software safety will also be discussed.

Knight, John C.↗

NASA trend analysis procedures

This publication is primarily intended for use by NASA personnel engaged in managing or implementing trend analysis programs. 'Trend analysis' refers to the observation of current activity in the context of the past in order to infer the expected level of future activity. NASA trend analysis was divided into 5 categories: problem, performance, supportability, programmatic, and reliability. Problem trend analysis uncovers multiple occurrences of historical hardware or software problems or failures in order to focus future corrective action. Performance trend analysis observes changing levels of real-time or historical flight vehicle performance parameters such as temperatures, pressures, and flow rates as compared to specification or 'safe' limits. Supportability trend analysis assesses the adequacy of the spaceflight logistics system; example indicators are repair-turn-around time and parts stockage levels. Programmatic trend analysis uses quantitative indicators to evaluate the 'health' of NASA programs of all types. Finally, reliability trend analysis attempts to evaluate the growth of system reliability based on a decreasing rate of occurrence of hardware problems over time. Procedures for conducting all five types of trend analysis are provided in this publication, prepared through the joint efforts of the NASA Trend Analysis Working Group.

Source record↗

Evaluation of Fieldbus and OPC for Advanced Life Support

FOUNDATION(Tm) Fieldbus and OP(TM) (OLE(TM)for Process Control) technologies were integrated into an existing control system for a crop growth chamber at NASA Ames Research Center. FOUNDATION(TM) Fieldbus is a digital, bi-directional, multi-drop, serial communications network which functions essentially as a LAN for sensors. FOUNDATION(TM) Fieldbus is heterarchical, with publishers and subscribers of data performing complex control functions at low levels without centralized control and its associated overhead. OPC(TM) is a set of interfaces which replace proprietary drivers with a transparent means of exchanging data between the fieldbus and applications. The objectives were: (1) to integrate FOUNDATION(TM) Fieldbus into existing ALS hardware and determine its overall effectiveness and reliability and, (2) to quantify any savings produced by using fieldbus and OPC technologies. We encountered several problems with the FOUNDATION(TM) Fieldbus hardware chosen. Our hardware exposed 100 data for each channel of the fieldbus. The fieldbus configurator software used to program the fieldbus was simply not adequate. The fieldbus was also not inherently reliable. It lost its settings twice during our tests for unknown reasons. OPC also had issues. It did not function at all as supplied, requiring substitution of some of its components with those from other vendors. It would stop working after a fixed period of time. Certain database calls eventually lock the machine. Overall, we would not recommend FOUNDATION(TM) Fieldbus: it was too difficult to implement with little overall added value. It also seems unlikely that FOUNDATION(TM) Fieldbus will gain sufficient penetration into the laboratory instrument market to ever be cost effective for the ALS community. OPC had good reliability and performance once a stable installation was achieved. It allowed a rapid change to an alternative software strategy when our first strategy failed. It is a cost effective solution to distributed control systems development.

Boulanger, Richard P.↗

Using magnetic tape technology for data migration

Magnetic tape and optical disk library units (jukeboxes) are satisfying the demand for high-capacity cost-effective storage. The choice between optical disk and magnetic tape technology must take into account the cost limitations as well as the performance and reliability requirements of the user environment. Library units require data management software in order to function in an automated and user-transparent way. The most common data management applications are backup and recovery, data migration, and archiving. The medium access patterns that these applications create will be described. Since the most user visible application is data migration, a queue simulator was developed to model its performance against a variety of library units. The major subject of this paper is the design and implementation of this simulator as well as some simulation results. The relative cost and reliability of magnetic tape versus optical disk library units is presented for completeness.

Therrien, David↗