Search NASA⌕ Search

SEARCH · Search NASA

Results for “adaptive stress testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Adaptive Stress Testing of Trajectory Predictions in Flight Management Systems

To find failure events and their likelihoods in flight-critical systems, we investigate the use of an advanced black-box stress testing approach called adaptive stress testing. We analyze a trajectory predictor from a developmental commercial flight management system which takes as input a collection of lateral waypoints and en-route environmental conditions. Our aim is to search for failure events relating to inconsistencies in the predicted lateral trajectories. The intention of this work is to find likely failures and report them back to the developers so they can address and potentially resolve shortcomings of the system before deployment. To improve search performance, this work extends the adaptive stress testing formulation to be applied more generally to sequential decision-making problems with episodic reward by collecting the state transitions during the search and evaluating at the end of the simulated rollout. We use a modified Monte Carlo tree search algorithm with progressive widening as our adversarial reinforcement learner. The performance is compared to direct Monte Carlo simulations and to the cross-entropy method as an alternative importance sampling baseline. The goal is to find potential problems otherwise not found by traditional requirements-based testing. Results indicate that our adaptive stress testing approach finds more failures and finds failures with higher likelihood relative to the baseline approaches.

adaptive stress testing↗

Adaptive Stress Testing: Finding Likely Failure Events with Reinforcement Learning

Finding the most likely path to a set of failure states is important to the analysis of safety-critical systems that operate over a sequence of time steps, such as aircraft collision avoidance systems and autonomous cars. In many applications such as autonomous driving, failures cannot be completely eliminated due to the complex stochastic environment in which the system operates.As a result, safety validation is not only concerned about whether a failure can occur, but also discovering which failures are most likely to occur. This article presents adaptive stress testing (AST), a framework for finding the most likely path to a failure event in simulation. We consider a general black box setting for partially observable and continuous-valued systems operating in an environment with stochastic disturbances. We formulate the problem as a Markov decision process and use reinforcement learning to optimize it. The approach is simulation-based and does not require internal knowledge of the system, making it suitable for black-box testing of large systems. We present different formulations depending on whether the state is fully observable or partially observable. In the latter case, we present a modified Monte Carlo tree search algorithm that only requires access to the pseudorandom number generator of the simulator to overcome partial observability. We also present an extension of the framework, called differential adaptive stress testing (DAST), that can find failures that occur in one system but not in another. This type of differential analysis is useful in applications such as regression testing, where we are concerned with finding areas of relative weakness compared to a baseline. We demonstrate the effectiveness of the approach on an aircraft collision avoidance application, where a prototype aircraft collision avoidance system is stress tested to find the most likely scenarios of near mid-air collision.

Verification and Validation↗

Differential Adaptive Stress Testing of Airborne Collision Avoidance Systems

The next-generation Airborne Collision Avoidance System (ACAS X) is currently being developed and tested to replace the Traffic Alert and Collision Avoidance System (TCAS) as the next international standard for collision avoidance. To validate the safety of the system, stress testing in simulation is one of several approaches for analyzing near mid-air collisions (NMACs). Understanding how NMACs can occur is important for characterizing risk and informingdevelopment of the system. Recently, adaptive stress testing (AST) has been proposed as a way to find the most likely path to a failure event. The simulation-based approach accelerates search by formulating stress testing as a sequential decision process then optimizing it using reinforcement learning. The approach has been successfully applied to stress test a prototype of ACAS Xin various simulated aircraft encounters. In some applications, we are not as interestedin the system's absolute performance as its performance relative to another system. Such situations arise, for example, during regression testing or when deciding whether a new system should replace an existing system. In our collision avoidance application, we are interested in finding cases where ACAS X fails but TCAS succeeds in resolving a conflict. Existing approaches do not provide an efficient means to perform this type of analysis. This paper extends the AST approach to differential analysis by searching two simulators simultaneously and maximizing the difference between their outcomes. We call this approach differential adaptive stress testing (DAST). We apply DAST to compare a prototype of ACAS X against TCAS and show examples of encounters found by the algorithm.

Lee, Ritchie↗

Validation of Image-Based Neural Network Controllersthrough Adaptive Stress Testing

Neural networks have become state-of-the-art for computer vision problems because of their ability to efficiently model complex functions from large amounts of data. While neural networks can be shown to perform well empirically fora variety of tasks, their performance is difficult to guarantee.Neural network verification tools have been developed that can certify robustness with respect to a given input image; however,for neural network systems used in closed-loop controllers,robustness with respect to individual images does not address multi-step properties of the neural network controller and itsenvironment. Furthermore, neural network systems interacting in the physical world and using natural images are operating in a black-box environment, making formal verification in-tractable. This work combines the adaptive stress testing (AST)framework with neural network verification tools to search for the most likely sequence of image disturbances that cause the neural network controlled system to reach a failure. Anautonomous aircraft taxi application is presented, and results show that the AST method finds failures with more likely image disturbances than baseline methods. Further analysis of AST results revealed an explainable cause of the failure, giving insight into the problematic scenarios that should be addressed.

Adaptive Stress Testing, Marabou, Deep Neural Netw↗

Adaptive Stress Testing of Collision Avoidance Systems for Small UASs with Deep Reinforcement Learning

The next-generation Airborne Collision Avoidance System for smaller UASs (ACAS sXu) is currently being developed and tested by the Federal Aviation Administration (FAA) to provide detect-and-avoid capability for small unmanned aircraft operating beyond line-of-sight. Due to the complexity and safety-critical nature of the system, safety validation is important not only for the certification of the final system, but also for informing changes during the iterative development process. In this paper, we analyze a prototype of ACAS sXu in simulated aircraft encounters to discover scenarios of small near mid-air collisions (sNMACs), an important safety event in which two aircraft come closer than 50 feet horizontally and 15 feet vertically. Due to the size and complexity of the system as well as rarity of sNMAC events, traditional methods such as Monte Carlo testing often require informed setup and targeting to elicit failures. However, such a dependence on domain knowledge can be incompatible with the independent verification and validation (IV&V) process, the aim of which is to discover unforeseen issues. To address these challenges, we apply an accelerated validation method called adaptive stress testing (AST) to find the most likely sNMAC scenarios without reliance on system introspection. AST uses reinforcement learning to adapt the search towards the most promising areas of the search space as it progresses. We use a state-of-the-art deep reinforcement learning algorithm, proximate policy optimization, to more efficiently search the large and continuous state space. We find that this approach significantly improves the performance of AST compared to a prior approach based on Monte Carlo tree search. We perform experiments using AST to find sNMAC events under various encounter configurations, varying parameters pertaining to dynamics and coordination. Our experiments show AST to be very effective at finding sNMAC scenarios. We summarize our findings, presenting high-level categories of discovered sNMACs and specific examples of encounters in each category.

aircraft collision avoidance↗

Adaptive Stress Testing: Using Reinforcement Learning to Find Failures in Safety-Critical Systems

Emerging applications in artificial intelligence, such as driverless cars and autonomous aircraft promise to be more efficient, cheaper to operate, and always available. However, ensuring the safety of these systems remains a major challenge to their certification and adoption. These autonomous systems are expected to routinely make safety-critical decisions where failures can have serious consequences including loss of life and property. Testing and validation techniques aim to identify and diagnose potential failures before the system is deployed. However, finding failure scenarios in autonomous systems can be very challenging due to high-dimensional and continuous state spaces, interaction with large environments over many time steps, and the rarity of failures. This talk presents Adaptive Stress Testing (AST), a simulation-based testing framework for finding the most likely path to a failure event of a safety-critical system. The key idea of AST is that stress testing can be formulated as a Partially Observable Markov Decision Process (POMDP), which enables reinforcement learning techniques to be used for finding failure events. Reinforcement learning algorithms can efficiently explore the search space and have been shown to scale to very large systems. We present applications of AST to find failures in various safety-critical systems including the aircraft collision avoidance systems, autonomous cars, and small unmanned aerial vehicles.

autonomous vehicles↗

Certification Considerations for Adaptive Stress Testing of Airborne Software

eduAdaptive Stress Testing (AST) has shown promise in identifying errant corner cases in complex software used in aerospace applications including Flight Management Systems (FMS). The strength of AST is performing test-based verification of complex aerospace software intensive systems at scale in simulated operational environments.Simulating and capturing the realistic operational complexities in integrated verification environments may exposeflaws in the softwareprior to field deployment, whereas the software may perform just fine to traditional requirements-basedunit and component level testing.AST can be used to test the whole system.Individual components may behave safely, but together can result in complex interactions and emergent failures, so it is important to test at the integrated system level.Motivated by the observed benefitsat the prototype proof of concept scale, this paper considers how AST may be integrated into a production workflow and used to generate objective evidence in a processthat delivers certified aerospace software.The research includes evaluation of alignment with both DO-178C and Overarching Properties(OP). The paper addresses questions such as “where should AST fit in the Plan for Software Aspects of Certification (PSAC) and Software Verification Plan (SVP), what aspects of AST do not fit, and what objectives does it satisfy?” The paper concludes that AST is in fact useful at locating errors in complex airborne application software and in doing so provides benefits to suppliers and end users. Furthermore, AST appears appropriate to add value in both DO-178Cbased and Overarching Properties based certification approaches.

certification↗

Adaptive Stress Testing of Airborne Collision Avoidance Systems

This paper presents a scalable method to efficiently search for the most likely state trajectory leading to an event given only a simulator of a system. Our approach uses a reinforcement learning formulation and solves it using Monte Carlo Tree Search (MCTS). The approach places very few requirements on the underlying system, requiring only that the simulator provide some basic controls, the ability to evaluate certain conditions, and a mechanism to control the stochasticity in the system. Access to the system state is not required, allowing the method to support systems with hidden state. The method is applied to stress test a prototype aircraft collision avoidance system to identify trajectories that are likely to lead to near mid-air collisions. We present results for both single and multi-threat encounters and discuss their relevance. Compared with direct Monte Carlo search, this MCTS method performs significantly better both in finding events and in maximizing their likelihood.

Verification and Validation↗

AdaStress

This is a tutorial on AdaStress, a tool for finding and analyzing the likeliest failures in a simulated system under test. The presentation outlines the adaptive stress testing framework, provides a demonstration of use, and showcases several examples of failure detection in a complex real-world system.

Reinforcement learning↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

reinforcement learning↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

Validation↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

Validation↗

Stress Testing California's Hydroclimatic Whiplash: Potential Challenges, Trade‐Offs and Adaptations in Water Management and Hydropower Generation

Abstract Inter‐annual precipitation in California is highly variable, and future projections indicate an increase in the intensity and frequency of hydroclimatic “whiplash.” Understanding the implications of these shocks on California's water system and its degree of resiliency is critical from a planning perspective. Therefore, we quantify the resilience of reservoir services provided by water and hydropower systems in four basins in the western Sierra Nevada. Using downscaled runoff from 10 climate model outputs, we generated 200 synthetic hydrologic whiplash sequences of alternating dry and wet years to represent a wide range of extremes and transitional conditions used as inputs to a water system simulation model. Sequences were derived from upper (wet) and lower (dry) quintiles of future streamflow projections (2030–2060). Results show that carryover storage was negatively affected in all basins, particularly in those with lower storage capacity. All basins experienced negative impacts on hydropower generation, with losses ranging from 5% to nearly 90%. Reservoir sizes and inflexible operating rules are a particular challenge for flood control, as in extremely wet years spillage averaged nearly the annual basins' total discharge. The reliability of environmental flows and agricultural deliveries varied depending on the basin, intensity, and duration of whiplash sequences. Overall, wet years temporarily rebound negative drought effects, and greater storage capacity results in higher reliability and resiliency, and lesser volatility in services. We highlight potential policy changes to improve flexibility, increase resilience, and better equip managers to face challenges posed by whiplash while meeting human and environmental needs.

Environmental Sciences & Ecology↗

Identifying Robust Decarbonization Pathways for the Western U.S. Electric Power System Under Deep Climate Uncertainty

Climate change threatens the resource adequacy of future power systems. Existing research and practice lack frameworks for identifying decarbonization pathways that are robust to climate-related uncertainty. We create such an analytical framework, then use it to assess the robustness of alternative pathways to achieving 60% emissions reductions from 2022 levels by 2040 for the Western U.S. power system. Our framework integrates power system planning and resource adequacy models with 100 climate realizations from a large climate ensemble. Climate realizations drive electricity demand; thermal plant availability; and wind, solar, and hydropower generation. Among five initial decarbonization pathways, all exhibit modest to significant resource adequacy failures under climate realizations in 2040, but certain pathways experience significantly less resource adequacy failures at little additional cost relative to other pathways. By identifying and planning for an extreme climate realization that drives the largest resource adequacy failures across our pathways, we produce a new decarbonization pathway that has no resource adequacy failures under any climate realizations. This new pathway is roughly 5% more expensive than other pathways due to greater capacity investment, and shifts investment from wind to solar and natural gas generators. Our analysis suggests modest increases in investment costs can add significant robustness against climate change in decarbonizing power systems. Our framework can help power system planners adapt to climate change by stress testing future plans to potential climate realizations, and offers a unique bridge between energy system and climate modeling.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Seventeen Ketogenic steroids excretion in crewmen in a 90-day manned test of an advanced regenerative life support system

Seventeen KGS (17-Ketogenic steroids) and Na/K were determined in 19 urine specimens collected by each of 4 crewmen during the 90-day test. The specimens represented 10% aliquots of 24-hour collections stored frozen onboard the simulator until passthrough. Electrolytes were analyzed immediately after sample passthrough while the steroids were determined post-test on aliquots of the original sample held at 203 K. Steroid data was corrected for body weight and also for analytical variation in the laboratory urine pool control. Long-term nonspecific responses to low-level stress appear to be reflected by the individual and group mean 17-KGS excretion patterns. The first 39 and the last 20 days of the test were significantly different--and presumably more stressful to the crew--than the period from days 39 to 67. Reduction of adrenocortical function during the mid-test phase is attributed to either an adaptation to chronic or intermittent stress or was the result of an actual reduction in the operational demands of the test during this time. Most remarkable of the metabolic findings is the prevalance of high Na/K ratios and an abrupt peak on day 74 for all 4 crewmen.

Myers, D. J.↗

Summary of Activities for Health Monitoring of Composite Overwrapped Pressure Vessels

This three-year project (FY12-14) will design and demonstrate the ability of new Magnetic Stress Gages for the measurement of stresses on the inner diameter of a Composite Overwrapped Pressure Vessel overwrap. The sensors are being tested at White Sands Testing Facility (WSTF) where the results will be correlated with a known nondestructive technique acoustic emission. The gages will be produced utilizing Meandering Winding Magnetometer (MWM) and/or MWM array eddy current technology. The ultimate goal is to utilize this technology for the health monitoring of Composite Overwrapped Pressure Vessels for all future flight programs. The first full-scale pressurization test was performed at WSTF in June 2012. The goals of this test were to determine adaptations of the magnetic stress gauge instrumentation that would be necessary to allow multiple sensors to monitor the vessel's condition simultaneously and to determine how the sensor response changes with sensor selection and orientation. The second full scale pressurization test was performed at WSTF in August 2012. The goals of this test were to monitor the vessel's condition with multiple sensors simultaneously, to determine the viability of the multiplexing units (MUX) for the application, and to determine if the sensor responses in different orientations are repeatable. For both sets of tests the vessel was pressured up to 6,000 psi to simulate maximum operating pressure. Acoustic events were observed during the first pressurization cycle. This suggested that the extended storage period prior to use of this bottle led to a relaxation of the residual stresses imparted during auto-frettage. The pressurization tests successfully demonstrated the use of multiplexers with multiple MWM arrays to monitor a vessel. It was discovered that depending upon the sensor orientation, the frequencies, and the sense element, the MWM arrays can provide a variety of complementary information about the composite overwrapped pressure vessel load conditions. For example, low frequency measurements can be used to monitor the overwrap thickness and changes associated with pressure level. High frequency data is dominated by the properties of the overwrap, including the fiber orientations and lay-up of the layers.

Russell, Rick↗

Test code for the assessment and improvement of Reynolds stress models

An existing two-dimensional, compressible flow, Navier-Stokes computer code, containing a full Reynolds stress turbulence model, was adapted for use as a test bed for assessing and improving turbulence models based on turbulence simulation experiments. To date, the results of using the code in comparison with simulated channel flow and over an oscillating flat plate have shown that the turbulence model used in the code needs improvement for these flows. It is also shown that direct simulation of turbulent flows over a range of Reynolds numbers are needed to guide subsequent improvement of turbulence models.

Rubesin, M. W.↗

Heart Rate Response During Mission-Critical Tasks After Space Flight

Adaptation to microgravity could impair crewmembers? ability to perform required tasks upon entry into a gravity environment, such as return to Earth, or during extraterrestrial exploration. Historically, data have been collected in a controlled testing environment, but it is unclear whether these physiologic measures result in changes in functional performance. NASA?s Functional Task Test (FTT) aims to investigate whether adaptation to microgravity increases physiologic stress and impairs performance during mission-critical tasks. PURPOSE: To determine whether the well-accepted postflight tachycardia observed during standard laboratory tests also would be observed during simulations of mission-critical tasks during and after recovery from short-duration spaceflight. METHODS: Five astronauts participated in the FTT 30 days before launch, on landing day, and 1, 6, and 30 days after landing. Mean heart rate (HR) was measured during 5 simulations of mission-critical tasks: rising from (1) a chair or (2) recumbent seated position followed by walking through an obstacle course (egress from a space vehicle), (3) translating graduated masses from one location to another (geological sample collection), (4) walking on a treadmill at 6.4 km/h (ambulation on planetary surface), and (5) climbing 40 steps on a passive treadmill ladder (ingress to lander). For tasks 1, 2, 3, and 5, astronauts were encouraged to complete the task as quickly as possible. Time to complete tasks and mean HR during each task were analyzed using repeated measures ANOVA and ANCOVA respectively, in which task duration was a covariate. RESULTS: Landing day HR was higher (P < 0.05) than preflight during the upright seat egress (7%+/-3), treadmill walk (13%+/-3) and ladder climb (10%+/-4), and HR remained elevated during the treadmill walk 1 day after landing. During tasks in which HR was not elevated on landing day, task duration was significantly greater on landing day (recumbent seat egress: 25%+/-14 and mass translation: 26%+/-12; P < 0.05). CONCLUSION: Elevated HR and increased task duration during postflight simulations of mission-critical tasks is suggestive of spaceflight-induced deconditioning. Following short-duration microgravity missions (< 16 d), work performance may be transiently impaired, but recovery is rapid.

Arzeno, Natalia M.↗