Search NASA⌕ Search

SEARCH · Search NASA

Results for “Support for Usability Evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Role of Trust and Usability in Enabling Spaceflight Crew Autonomy

Future long duration exploration missions will require an increased use of onboard automated systems as spaceflight crews venture further than before and have longer communications delays with ground support. The design of these systems must support appropriate crew trust and have sufficient usability to enable spaceflight crew autonomy or risk being misused while crews wait to communicate with ground support. We evaluated trust & usability in our self-scheduling tool, Playbook, for crew mission timelines. Data was collected in a controlled lab experiment with 31 participants. Participants in the study conducted two tasks: scheduling, where participants were responsible for scheduling a majority of a day's operational tasks, and rescheduling, where participants were provided a schedule and asked to reschedule higher priority activities. We found a significant correlation between system trust and usability, irrespective of self-scheduling tasks. We conclude that system usability may play a bigger role in how trust is learned while conducting novel crew autonomy tasks such as self-scheduling. Future research should investigate the role of usability to encourage appropriate trust in onboard automated crew systems and enable crew autonomy.

crew autonomy↗

Automating CPM-GOMS

CPM-GOMS is a modeling method that combines the task decomposition of a GOMS analysis with a model of human resource usage at the level of cognitive, perceptual, and motor operations. CPM-GOMS models have made accurate predictions about skilled user behavior in routine tasks, but developing such models is tedious and error-prone. We describe a process for automatically generating CPM-GOMS models from a hierarchical task decomposition expressed in a cognitive modeling tool called Apex. Resource scheduling in Apex automates the difficult task of interleaving the cognitive, perceptual, and motor resources underlying common task operators (e.g. mouse move-and-click). Apex's UI automatically generates PERT charts, which allow modelers to visualize a model's complex parallel behavior. Because interleaving and visualization is now automated, it is feasible to construct arbitrarily long sequences of behavior. To demonstrate the process, we present a model of automated teller interactions in Apex and discuss implications for user modeling. available to model human users, the Goals, Operators, Methods, and Selection (GOMS) method [6, 21] has been the most widely used, providing accurate, often zero-parameter, predictions of the routine performance of skilled users in a wide range of procedural tasks [6, 13, 15, 27, 28]. GOMS is meant to model routine behavior. The user is assumed to have methods that apply sequences of operators and to achieve a goal. Selection rules are applied when there is more than one method to achieve a goal. Many routine tasks lend themselves well to such decomposition. Decomposition produces a representation of the task as a set of nested goal states that include an initial state and a final state. The iterative decomposition into goals and nested subgoals can terminate in primitives of any desired granularity, the choice of level of detail dependent on the predictions required. Although GOMS has proven useful in HCI, tools to support the construction of GOMS models have not yet come into general use.

GOMS↗

Development of a preprototype vapor compression distillation water recovery subsystem

The activities involved in the design, development, and test of a preprototype vapor compression distillation water recovery subsystem are described. This subsystem, part of a larger regenerative life support evaluation system, is designed to recover usable water from urine, urinal rinse water, and concentrated shower and laundry brine collected from three space vehicle crewmen for a period of 180 days without resupply. Details of preliminary design and testing as well as component developments are included. Trade studies, considerations leading to concept selections, problems encountered, and test data are also presented. The rework of existing hardware, subsystem development including computer programs, assembly verification, and comprehensive baseline test results are discussed.

Johnson, K. L.↗

Wearable Biosensor Monitor to Support Autonomous Crew Health and Performance

During future human exploration missions, spaceflight crews will encounter adverse health outcomes and decrements in performance during the missions and for long term health. The Psychophysiological Research Lab is currently conducting a technology demonstration of a prototype wearable biosensor system (Astroskin) on astronaut surrogates during 30 day missions within NASAs Human Exploration Research Analog (HERA). HERA, located at Johnson Spaceflight Center, represents a flight analog for simulation of isolation, confinement and remote conditions of mission exploration scenarios. Astroskin, an autonomous medical monitoring system, consists of an intelligent garment for the upper body and a headband fitted with sensors, and associated software and hardware that can measure, transmit and store vital signs (ECG, respiration, blood pressure, etc.), sleep quality and activity level of the wearer. One of the main goals of the technology demonstration is to evaluate the performance of the Astroskin system (i.e., data quality), crew usability and comfort. To validate the Astroskin as a viable option for use during future spaceflight missions, crew data from HERA are sent to Ames researchers for processing and analyses. Crew surveys on usability and comfort will also be evaluated. This work supports the Exploration Medical Capabilities element within NASAs Human Research Program.

space analog↗

Mission Evaluation Room Intelligent Diagnostic and Analysis System (MIDAS)

The role of Mission Evaluation Room (MER) engineers is to provide engineering support during Space Shuttle missions, for Space Shuttle systems. These engineers are concerned with ensuring that the systems for which they are responsible function reliably, and as intended. The MER is a central facility from which engineers may work, in fulfilling this obligation. Engineers participate in real-time monitoring of shuttle telemetry data and provide a variety of analyses associated with the operation of the shuttle. The Johnson Space Center's Automation and Robotics Division is working to transfer advances in intelligent systems technology to NASA's operational environment. Specifically, the MER Intelligent Diagnostic and Analysis System (MIDAS) project provides MER engineers with software to assist them with monitoring, filtering and analyzing Shuttle telemetry data, during and after Shuttle missions. MIDAS off-loads to computers and software, the tasks of data gathering, filtering, and analysis, and provides the engineers with information which is in a more concise and usable form needed to support decision making and engineering evaluation. Engineers are then able to concentrate on more difficult problems as they arise. This paper describes some, but not all of the applications that have been developed for MER engineers, under the MIDAS Project. The sampling described herewith was selected to show the range of tasks that engineers must perform for mission support, and to show the various levels of automation that have been applied to assist their efforts.

Pack, Ginger L.↗

Usability Evaluation of Spot and Runway Departure Advisor (SARDA) Concept in Dallas/Fort Worth Airport Tower Simulation

Spot and Runway Departure Advisor (SARDA) is a proposed decision-support tool for air traffic control tower controllers for reducing taxi delay and optimizing the departure sequence. In the present study, the tool's usability was evaluated to ensure that its claimed performance benefits are not being realized at the cost of increasing the work burden on controllers. For the evaluation, workload ratings and questionnaire responses collected during a human-in-the-loop simulation experiment were analyzed to assess the SARDA advisories' effects on the controllers' ratings on cognitive resources (e.g., workload, spare attention) and satisfaction. The results showed that SARDA reduced the controllers' workload and increased their spare attention. It also made workload and attention levels less susceptible to the effects of increases in the traffic load. The questionnaire responses suggested that the controllers generally were satisfied with the ease of use of the tool and the objectives of the SARDA concept, but with some caution. To gain more trust from controllers, the the reasoning behind advisories may need to be made more transparent to them.

human factors↗

Evaluation of Self-Scheduling Exercises Completed by Analog Crewmembers in NASA's Human Exploration Research Analog (HERA)

NASA human spaceflight missions are inherently dynamic and require frequent scheduling changes in order to adapt to changing mission priorities and objectives. Tactical level changes to the mission plan are traditionally made by a team of expert planners and operations specialists on the ground. However, astronauts are expected to execute missions more autonomously during future long duration missions. Astronauts will need to take on some of the responsibility of managing their own schedule while still abiding by the numerous constraints required by human spaceflight operations. This paper summarizes salient elements of crew performance in NASA’s Human Exploration Research Analog Campaign 3. Analog crewmembers completed a series of self-scheduling exercises to evaluate Playbook’s usability towards enabling self-scheduling without support from ground control. Playbook is a self-scheduling software tool designed and developed by our team. We also investigated how to best communicate self-scheduling tasks and constraints to the crew in order to facilitate efficient self-scheduling during isolation in a realistic environment. Our analysis identified that 30 minutes was sufficient to complete complex self-scheduling tasks. Our evaluation also identified differences between individual and collaborative performance; analog crewmembers completed self-scheduling exercises more quickly as a team as opposed to individually and reported lower subjective difficulty ratings overall.

HERA↗

Evaluation of Usability and Workload Associated with Paper Strips as Compared to Virtual Flight Strips Used for Ramp Operations

This paper describes a study comparing the use of paper strips with virtual flight strips depicted on a new user interface, the Ramp Traffic Console (RTC), designed for use by ramp controllers to be used in place of paper strips. A Human-In-the-Loop (HITL) experiment was performed as the fifth in a series of six HITL simulation studies designed to evaluate a pushback Decision Support Tool (DST) concept for Charlotte Douglas International Airport (CLT). Workload and usability were assessed in post-run and post-study questionnaires. In the RTC virtual flight strip condition, post-run questionnaire results show lower workload ratings across all aspects of workload; additionally a trend is found toward increased usability ratings. Post-study questionnaire results indicate a preference for RTC over paper strips. Additional research is suggested with more training runs and a greater number of participants to increase statistical power. It is also suggested that this new technology be re-evaluated as a part of the ATD-2 (Airspace Technology Demonstration 2) field testing activities.

Human Factors↗

Assessment of Electric Grid Transmission System Simulator for Human Factors Research

Over the past decades, various technologies have been developed for electric grid operations to support clean energy, meet rising electricity demands, and address infrastructure concerns. However, the human factors aspect is often overlooked during rapid integration. Questions persist about how these technologies impact human performance. Simulators play a critical role in supporting investigation of human factors design concepts and conducting comprehensive usability testing to evaluate human performance and assess human reliability. This paper aims to address human factors research simulator requirements and conduct a comparative study of six different simulators. A detailed evaluation reveals that the evaluated simulators lack the ability to customize user interfaces. Additionally, their user interface designs do not fulfill the basic human factors design principles, potentially leading to increased response variability and reduced statistical power when conducting controlled experimental research. In the future, it is essential to develop scripting tools to integrate customizable user interfaces and simulation models, ensuring meeting research requirements.

Li, Ruixuan↗

Cross-Measurement Comparisons for a CFD Validation Dataset on Mach 2.5 Axisymmetric Turbulent Shock-Wave/Boundary-Layer Interactions

Experimental data for a shock-wave/boundary-layer interaction has been collected using multiple measurement techniques. Unfortunately, diversity of methods for acquisition begets an aggregation of data which does not directly quantify the same properties of the flowfield. The objective in this paper was to (a) present the collection of flowfield measurements in one venue in a format more usable for those validating CFD models against the data, (b) evaluate the degree to which the experimental data support each other, and (c) highlight the differences and relative advantages/shortcomings of each measurement techniques. To present the various measured quantities into common format, CFD results from a companion paper also submitted for presentation at this meeting are utilized. For the case where the boundary layer remains attached, there is agreement between the various measurements as well as with Reynolds-averaged Navier-Stokes simulation solutions. As the impinging shock strength is increased beyond the point of separating the boundary layer, the congruity of the data wanes. Generally, agreement among the measurements exceeds the degree to which the CFD solutions agree with experiments. This suggests that unmodeled physical phenomena, such as transient motion of the reflected shock and separation bubble, give rise to the discrepancies observed in the computational results.

Shock-wave Boundary-layer interaction↗

Detecting Process Equipment Failures Using Acoustic Data and Machine Learning

Nuclear power plant (NPP) process equipment such as fans, motors, valves, and pumps generate frequent or continuous noise, and deviations from the normal operational sounds made by this equipment can indicate potential issues. These deviations can be identified via automated acoustic anomaly detection, which involves using acoustic sensors (i.e., microphones) alongside detection algorithms to continuously monitor for changes in acoustic signatures. This task is made challenging by the substantial background noise that exists, such as operators opening and closing doors, manipulating valves, and conversing—in addition to typical plant noises. In collaboration with a nuclear power utility partner, this effort assessed the efficacy of acoustic anomaly detection when using a specific acoustic sensor that compresses data into a fixed set of features that are transferable over a standard Internet of Things communication protocol, thereby improving usability but potentially degrading detection performance. Two methods of performing automated acoustic anomaly detection were evaluated: one-class support vector machine (OC-SVM) and isolation forest (iForest). To enable the use of high-quality acoustic data encompassing both normal and anomalous conditions, the study utilized the publicly available Malfunctioning Industrial Machine Investigation and Inspection dataset, which includes real measured acoustic sensor data for a range of equipment types, model numbers, and signal-to-noise ratios (SNRs), along with a benchmark set of detection results. Using this dataset, the methods were tested and then compared against the benchmark results. The results indicated that although the specific acoustic sensor did not enable as rich a feature set extraction, the proposed methods with the limited feature set performed just as well. This provides solid justification for both the methods and the use of the proposed acoustic sensor.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Playbook for UAS: UX of Goal-Oriented Planning & Execution

We are evaluating the tool Playbook, a mission planning and scheduling tool used for analog missions and research experiments at NASA. We are adapting Playbook to support an unmanned aerial vehicle (UAV) demonstration that will simulate a "search and recover" mission. To support this demonstration, we need to test this software and its usability with human subjects. Playbook will be used to plan, edit, and monitor simulated UAV swarms, and we are interested in evaluating these capabilities. Allocation of roles and responsibilities between human-automation systems is key to promoting productive cooperation between these two agents. When guiding swarms the user should not be concerned with moving waypoints, adjusting heading and altitude or any other small actions that traditional pilots control. Providing situational awareness through sensor data including radar, heat, live camera feed, and topographic information are key when managing multiple UAVs together. First responders need a good picture of disaster areas in order to make informed decisions. These sensors can initiate movements in the mission that indicate when the human should take control of single aircraft.

UAV swarms↗

Flight Deck Design of a Hybrid Turbine/Electric Passenger Aircraft

NASA is exploring the development of a 180-passenger subsonic single engine aft turbine aircraft, The aft turbine provides electric power in a hybrid design to wing mounted electric engines, creating a highly efficient, high-bypass-ratio fan equivalent. The SUbsonic Single Aft eNgine (SUSAN) aircraft is being developed as a sustainable subsonic regional aircraft that seeks to reduce emission levels by 50% in the next few decades. Pilot-in-the-loop studies were conducted at the NASA Langley Research Center in Hampton, Virginia, to explore the flight deck design for the hybrid electric aircraft. Following modern trends in commercial aircraft flight decks with full time augmented controls and a quiet and dark philosophy, single throttle and simplified engine displays were developed for the SUSAN aircraft. The aircraft includes a single aft mounted turbine engine and 16 wing mounted electric fans. The final design was developed from feedback received during an earlier pilot-in-the-loop study where one, two, and three throttles were tested in standard airline operations, including various failures of the turbine and electric engines. Current flight deck designs normally provide control inceptors for each propulsion engine and an engine display for all primary aircraft engine parameters. With full time augmentation expected, a single throttle control with autothrottle always engaged, even during failures, is desired. Augmentation of flight controls using distributed thrust also requires full time control of the electric engines using automation. Additionally, electric engine thrust is augmented during climb based on battery state of charge. Thrust augmentation changes faster than human reaction time and therefore requires full-time automation. A pilot-in-the-loop study was conducted at the NASA Langley Research Center in Hampton, Virginia, to test the final design of the single throttle with simplified engine displays. Fourteen airline pilots evaluated the single throttle and engine display concept. Electric engine failures included one, four symmetric, and eight non-symmetric electric engine failures. The turbine engine was evaluated for complete and partial failure during critical phases of flight to include takeoff as well as enroute. Unexpected go-arounds increase workload and require significant throttle manipulation. Go-arounds were included to ensure the single throttle was usable for all phases of flight. Failures during takeoff required a return to the departure field and failures enroute required a diversion except for one and four electric engine failures as these failures did not affect aircraft flyability or range. There are currently no Part 25 aircraft certified with hybrid systems or electric engines with batteries as emergency propulsion. For turbine engine failures in the SUSAN aircraft design, range is limited to 30 minutes at full power. Battery state of charge and battery health displays were developed and tested for usability and to determine how well they supported pilot decisions for alternate airports during emergency diversions. Novel displays using shape and color were developed to provide immediate feedback when state of charge became critical. This paper details the pilot study including pilot feedback supporting the potential for increased automation and a single throttle control. Detailed recommendations are provided for a novel single throttle control and additional pilot controls to support selection of engines during start, shutdown, and engine troubleshooting procedures. This design deviates significantly from current practice of providing throttles for each propulsion engine. Engine display recommendations are provided based on pilot feedback during a guided post-evaluation interview. Battery state of charge and battery health display recommendations were collected from all airline crews. The simplified engine displays design was rated excellent as measured with a usability scale. Quantitative metrics include airspeed tracking, time to complete checklists, time to make diversion decisions and the quality of the diversion decision. Recommendations for future studies are documented with supporting research and current observations about upcoming flight deck certifications.

autothrottle↗

Flight Deck Design of a Hybrid Turbine/Electric Passenger Aircraft

NASA is exploring the development of a 180-passenger subsonic single engine aft turbine aircraft, The aft turbine provides electric power in a hybrid design to wing mounted electric engines, creating a highly efficient, high-bypass-ratio fan equivalent. The SUbsonic Single Aft eNgine (SUSAN) aircraft is being developed as a sustainable subsonic regional aircraft that seeks to reduce emission levels by 50% in the next few decades. Pilot-in-the-loop studies were conducted at the NASA Langley Research Center in Hampton, Virginia, to explore the flight deck design for the hybrid electric aircraft. Following modern trends in commercial aircraft flight decks with full time augmented controls and a quiet and dark philosophy, single throttle and simplified engine displays were developed for the SUSAN aircraft. The aircraft includes a single aft mounted turbine engine and 16 wing mounted electric fans. The final design was developed from feedback received during an earlier pilot-in-the-loop study where one, two, and three throttles were tested in standard airline operations, including various failures of the turbine and electric engines. Current flight deck designs normally provide control inceptors for each propulsion engine and an engine display for all primary aircraft engine parameters. With full time augmentation expected, a single throttle control with autothrottle always engaged, even during failures, is desired. Augmentation of flight controls using distributed thrust also requires full time control of the electric engines using automation. Additionally, electric engine thrust is augmented during climb based on battery state of charge. Thrust augmentation changes faster than human reaction time and therefore requires full-time automation. A pilot-in-the-loop study was conducted at the NASA Langley Research Center in Hampton, Virginia, to test the final design of the single throttle with simplified engine displays. Fourteen airline pilots evaluated the single throttle and engine display concept. Electric engine failures included one, four symmetric, and eight non-symmetric electric engine failures. The turbine engine was evaluated for complete and partial failure during critical phases of flight to include takeoff as well as enroute. Unexpected go-arounds increase workload and require significant throttle manipulation. Go-arounds were included to ensure the single throttle was usable for all phases of flight. Failures during takeoff required a return to the departure field and failures enroute required a diversion except for one and four electric engine failures as these failures did not affect aircraft flyability or range. There are currently no Part 25 aircraft certified with hybrid systems or electric engines with batteries as emergency propulsion. For turbine engine failures in the SUSAN aircraft design, range is limited to 30 minutes at full power. Battery state of charge and battery health displays were developed and tested for usability and to determine how well they supported pilot decisions for alternate airports during emergency diversions. Novel displays using shape and color were developed to provide immediate feedback when state of charge became critical. This paper details the pilot study including pilot feedback supporting the potential for increased automation and a single throttle control. Detailed recommendations are provided for a novel single throttle control and additional pilot controls to support selection of engines during start, shutdown, and engine troubleshooting procedures. This design deviates significantly from current practice of providing throttles for each propulsion engine. Engine display recommendations are provided based on pilot feedback during a guided post-evaluation interview. Battery state of charge and battery health display recommendations were collected from all airline crews. The simplified engine displays design was rated excellent as measured with a usability scale. Quantitative metrics include airspeed tracking, time to complete checklists, time to make diversion decisions and the quality of the diversion decision. Recommendations for future studies are documented with supporting research and current observations about upcoming flight deck certifications.

autothrottle↗

Usability Evaluation Technical Report: A Usability Evaluation of the Communication Tool Prototype Prepared in Adobe XD for High Density Vertiplex (HDV) Team

Purpose: As cities get more crowded, the roadway infrastructure cannot keep up with the travel demands. Aviation can be a solution. Organizations supporting NASA’s Urban Air Mobility (UAM) concept are conducting studies on feasible concepts of operations for the new air traffic management system required to implement UAM. NASA’s High Density Vertiplex (HDV) team conducts studies using both live flights and simulations. Significance: The goal of this study is to evaluate the effectiveness of a communication tool developed for this team to enhance team situation awareness. Specifically, the study aims to determine if the communication tool is effective in allowing NASA-Langley and NASA-Ames team members playing different roles to understand what others are doing so that completed steps can be communicated to the whole team. Methods: This study asks participants to conduct four distinct scenarios using an Adobe XD prototype. Participants included five members of the target audience, who each received up to 15 minutes of training on how to use the Adobe XD prototype. The participants were NASA-Ames or NASA-Langley researchers, developers or tech leads who had 0 to 15 years of experience. Their ages ranged from 20 to 40 years old, and they were of both genders. After training, each participant was asked to do 4 tasks of varying difficulty, using the prototype testbed. After each task they reported on task difficulty and after all tasks were completed they were asked to fill out the System Usability Survey (SUS). Performance will be assessed through task completion time and accuracy rating. Results: User perception was evaluated by the SUS which showed a mean score over the six participants of 84. All the participants were above the acceptable range which was 70 and nobody was in the marginal usability zone between 50 and 70. In addition, most of the participants did well in terms of time on task, success rate and task difficulty on all the tasks except for Task 3. Participants had two minutes to do the task correctly to succeed at the task, otherwise it was marked as a failure. The geomean for Task 1 for completion time was 22.20; Task 2 was 58.28; Task 3 was 84.68 and Task 4 was 19.61. In Task 3, three out of six people failed this task but all five scored this with an average score 4.2 out of 5 for self-rating on task difficulty.

human factors↗

Preliminary Design of an 'Autonomous Medical Response Agent' Interface Prototype for Long Duration Spaceflight

Major challenges for astronauts in future long-duration exploration missions (LDEMs) will be that crewmembers are not expected to be medical professionals, may be under high workload and stress, are facing physiological challenges caused by spaceflight, and will have limited, delayed voice communications with medical support from Earth. An autonomous medical response agent (AMRA) is envisioned to help astronauts address medical complaints, develop a differential diagnosis, and guide self-treatment until a healthy state is restored. AMRA develops a process of personalized diagnosis and treatment through a Bayesian predictive control system that recommends therapeutic control actions including diagnostic tests and treatments to crewmembers (Menon, 2020). The Human Computer Interaction (HCI) lab from NASA Ames Research Center’s Human Systems Integration Division (Code TH) has collaborated with Nahlia Inc in human-centered design augmentation research for AMRA. The project, titled Design of ‘Autonomous Medical Response Agent Interface Prototype for Long Duration Spaceflight, has been funded by the Translational Research Institute for Space Health (TRISH) and introduces an interactive user-interface prototype that guides astronauts through self-diagnosis, treatment, and rehabilitation while communicating with remote specialists in ground support (most notably a patient’s flight surgeon). Our project develops the interaction design for the crewmember using AMRA through user research, iterative design, and usability testing to evaluate the user interface and workflow designed. The interface design deliverable for this project, titled AMRA Aggregate Information Display (AMRA AID) is an integrated information display system for comprehensive autonomous medical guidance, diagnosis, and treatment of in-flight medical conditions experienced by crewmembers. AMRA AID demonstrates how we might ensure crew autonomy, increase the crew’s medical capabilities, and decrease cognitive burden within a front-end user interface. AMRA AID refrains from relying on input from ground or mission control for self-treatment of medical issues—though ground awareness and communication with ground is maintained as a means of ensuring trust between mission control and crew. AMRA AID demonstrates how the crew’s on-board medical system might integrate with information from vehicle monitoring and crew schedule, without assuming causal relationships. AMRA AID’s comprehensive view enables efficient information access for both crew and ground support, reducing cognitive burden in the event of an unplanned or emergency medical incident and enabling informed analytical decisions to be made based on both crew and vehicle health. Human-centered design augmentation advanced within the prototype included: enhanced workflow and treatment guidance for two medical scenarios for a non-specialist user base with various levels of medical training, interaction design which considered speech (conversational user interface) elements and on-screen interactions to be developed in future iterations of the project, communication design and functional requirements relevant to self-care versus caring for another astronaut, as well as user testing of the prototype with an international space medical community. This project arrives at critical findings regarding usability needs, communication requirements, and integrated information requirements for a future technology interface functioning to increase confidence between ground support and LDEM crewmembers.

TRISH↗

Usability Evaluation of Indicators of Energy-Related Problems in Commercial Airline Flight Decks

A series of pilot-in-the-loop flight simulation studies were conducted at NASA Langley Research Center to evaluate indicators aimed at supporting the flight crew’s awareness of problems related to energy states. Indicators were evaluated utilizing state-of-the-art flight deck systems such as on commercial air transport aircraft. This paper presents results for four technologies: (1) conventional primary flight display speed cues, (2) an enhanced airspeed control indicator, (3) a synthetic vision baseline that provides a flight path vector, speed error, and an acceleration cue, and (4) an aural airspeed alert that triggers when current airspeed deviates beyond a specified threshold from the selected airspeed. Full-mission high-fidelity flight simulation studies were conducted using commercial airline crews. Crews were paired by airline for common crew resource management procedures and protocols. Scenarios spanned a range of complex conditions while emulating several causal factors reported in recent accidents involving loss of energy state awareness by pilots. Data collection included questionnaires administered at the completion of flight scenarios, aircraft state data, audio/video recordings of flight crew, eye tracking, pilot control inputs, and researcher observations. Questionnaire response data included subjective measures of workload, situation awareness, complexity, usability, and acceptability. This paper reports relevant findings derived from subjective measures as well as quantitative measures.

Evans, Emory T.↗

Heuristic Evaluation Methods Applied to a Predictive Maintenance Chatbot

The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER’s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use.

99 - GENERAL AND MISCELLANEOUS↗