Search NASA⌕ Search

SEARCH · Search NASA

Results for “task scenario”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Communication variations and aircrew performance

The relationship between communication variations and aircrew performance (high-error vs low-error performances) was investigated by analyzing the coded verbal transcripts derived from the videotape records of 18 two-person air transport crews who participated in a high-fidelity, full-mission flight simulation. The flight scenario included a task which involved abnormal operations and required the coordinated efforts of all crew members. It was found that the best-performing crews were characterized by nearly identical patterns of communication, whereas the midrange and poorer performing crews showed a great deal of heterogeneity in their speech patterns. Although some specific speech sequences can be interpreted as being more or less facilitative to the crew-coordination process, predictability appears to be the key ingredient for enhancing crew performance. Crews communicating in highly standard (hence predictable) ways were better able to coordinate their task, whereas crews characterized by multiple, nonstandard communication profiles were less effective in their performance.

Kanki, Barbara G.↗

Admittance model for the shuttle remote manipulator system in four configurations

A possible scenario for robot task performance in space is to mount two small, dexterous arms to the end of the Shuttle Remote Manipulator System (SRMS). As these small robots perform tasks, the flexibility of the SRMS may cause unsuccessful task executions. In order to simulate the dynamic coupling between the SRMS and the arms, admittance models of the SRMS in four brakes locked configurations were developed. The admittance model permits calculation of the SRMS end-effector response due to end-effector disturbing forces. The model will then be used in conjunction with a Stewart Platform, a vehicle emulation system. An application of the admittance model was shown by simulating the disturbing forces using two SRMS payloads, the Dextrous Orbital Servicing System (DOSS) manipulator and DOSS carrying a 1000 lb. cylinder. Mode by mode comparisons were conducted to determine the minimum number of modes required in the admittance model while retaining dynamic fidelity. It was determined that for all four SRMS configurations studied, between 4 and 6 modes of the SRMS structure (depending on the excitation loads) were sufficient to retain tolerance of 0.01 inches and 0.01 deg. These tolerances correspond to the DOSS manipulator carrying no object. When the DOSS carries the 1000 lb. cylinder, between 15 and 20 modes were sufficient, approximately three or four times as many modes as for the unloaded case.

Papadopoulos, Loukas↗

The Nitty Gritty: How We Make Analogs Work

NASA's Human Research Program (HRP) is becoming increasingly reliant on Isolated, Confined and Controlled (ICC) analogs to accomplish many of its research objectives. Compared to other research platforms, ICC analogs present a unique set of operational challenges that must be addressed in order to ensure a high fidelity research environment. In particular, the Human Exploration Research Analog (HERA) habitat, which is classified as an ICC environment, has been developed over the past three years to accommodate the operational needs of research investigations from each of the HRP Elements. During the development period, various types of requirements have contributed to the current operational model, which strives to achieve the highest possible level of mission fidelity with limited resources. This presentation will focus on the operational aspects of the HERA habitat, with emphasis on how we develop the analog research environment to meet researchers' needs. Specific discussion topics include mission scenario development, operational tasks, mission timeline integration, stressor implementation, console support, and improvements based on lessons learned. The information is intended to help investigators better understand the details behind HERA operations and the benefits to their research goals.

Self, A. L.↗

The space station assembly phase: Flight telerobotic servicer feasibility. Volume 2: Methodology and case study

A methodology is described for examining the feasibility of a Flight Telerobotic Servicer (FTS) using two assembly scenarios, defined at the EVA task level, for the 30 shuttle flights (beginning with MB-1) over a four-year period. Performing all EVA tasks by crew only is compared to a scenario in which crew EVA is augmented by FTS. A reference FTS concept is used as a technology baseline and life-cycle cost analysis is performed to highlight cost tradeoffs. The methodology, procedure, and data used to complete the analysis are documented in detail.

Smith, Jeffrey H.↗

The flight telerobotic servicer Tinman concept: System design drivers and task analysis

A study was conducted to develop a preliminary definition of the Flight Telerobotic Servicer (FTS) that could be used to understand the operational concepts and scenarios for the FTS. Called the Tinman, this design concept was also used to begin the process of establishing resources and interfaces for the FTS on Space Station Freedom, the National Space Transportation System shuttle orbiter, and the Orbital Maneuvering vehicle. Starting with an analysis of the requirements and task capabilities as stated in the Phase B study requirements document, the study identified eight major design drivers for the FTS. Each of these design drivers and their impacts on the Tinman design concept are described. Next, the planning that is currently underway for providing resources for the FTS on Space Station Freedom is discussed, including up to 2000 W of peak power, up to four color video channels, and command and data rates up to 500 kbps between the telerobot and the control station. Finally, an example is presented to show how the Tinman design concept was used to analyze task scenarios and explore the operational capabilities of the FTS. A structured methodology using a standard terminology consistent with the NASA/National Bureau of Standards Standard Reference Model for Telerobot Control System Architecture (NASREM) was developed for this analysis.

Andary, J. F.↗

Radiation Accidents and Malicious Events – Scenarios and Scope of the Work of ICRP Task Group 120

The International Commission on Radiological Protection (ICRP) Task Group 120 (TG120) is developing ICRP recommendations for radiological protection for a wide range of radiation accidents and malicious events, complementing those given in ICRP Publication 146 (2020) for large nuclear accidents. The scope includes accidents involving criticalities, operating faults, and fires and explosions in nuclear facilities, inadvertent damage to sealed radiation sources, as well as malicious events, such as sabotage of nuclear facilities or materials, use of radiological dispersal devices, the contamination of food and drinking water supplies, and the deployment of nuclear weapons. A template has been designed to collate relevant information on a wide range of case studies and hypothetical malicious scenarios to ensure that the recommendations developed are broadly applicable and comprehensive. For all scenarios, a graded approach to protection is being taken, accepting that specific guidance may be required for some distinctive aspects, for example, protection during times of armed conflict. This paper provides an overview of the scenarios and scope of the work of TG120, including some of the radiological and non-radiological impacts of radiation emergencies, along the response and recovery timeline.

ICRP↗

Determining Desirable Cursor Control Device Characteristics for NASA Exploration Missions

The Crew Exploration Vehicle (CEV) that will travel to the moon and Mars, and all future Exploration vehicles and habitats will be highly computerized, necessitating an accurate method of interaction with the computers. The design of a cursor control device will have to take into consideration g-forces, vibration, gloved operations, and the specific types of tasks to be performed. The study described here is being undertaken to begin identifying characteristics of cursor control devices that will work well for the unique Exploration mission environments. The objective of the study is not to identify a particular device, but to begin identifying design characteristics that are usable and desirable for space missions. Most cursor control devices have strengths and weaknesses; they are more appropriate for some tasks and less suitable for others. The purpose of this study is to collect some initial usability data on a large number of commercially available and proprietary cursor control devices. A software test battery was developed for this purpose. Once data has been collected using these low-level, basic point/click/drag tasks, higher fidelity, scenario-driven evaluations will be conducted with a reduced set of devices. The standard tasks used for testing cursor control devices are based on a model of human movement known as Fitts law. Fitts law predicts that the time to acquire a target is logarithmically related to the distance over the target size. To gather data for analysis with this law, fundamental, low-level tasks are used such as dragging or pointing at various targets of different sizes from various distances. The first four core tasks for the study were based on the ISO 9241-9:(2000) document from the International Organization for Standardization that contains the requirements for non-keyboard input devices. These include two pointing tasks, one dragging and one tracking task. The fifth task from ISO 9241-9, the circular tracking task was not used because it is a movement that is not applicable to most of the applications used on aviation displays. Additionally, we opted to add a multi-size and multi-distance pointing task, and two ecologically more valid tasks which included text selection, and interaction with drop down menus, sliders, and checkboxes. The Visual Basic test battery tracks the task and trial numbers, measures the pointing, tracking or dragging time, as well as the number and types of errors. The testing session includes a practice set for each input device, then the randomized 7 tasks, and finally a questionnaire about the device. This is repeated for all the devices tested within a session. The experiment is a within-subjects design, with participants returning for multiple sessions to test additional devices. The input devices will be compared based on objective performance data from the tasks, as well as subjective feedback and ratings on the questionnaire.

Aniko Sandor↗

The impact of physical and mental tasks on pilot mental workoad

Seven instrument-rated pilots with a wide range of backgrounds and experience levels flew four different scenarios on a fixed-base simulator. The Baseline scenario was the simplest of the four and had few mental and physical tasks. An activity scenario had many physical but few mental tasks. The Planning scenario had few physical and many mental taks. A Combined scenario had high mental and physical task loads. The magnitude of each pilot's altitude and airspeed deviations was measured, subjective workload ratings were recorded, and the degree of pilot compliance with assigned memory/planning tasks was noted. Mental and physical performance was a strong function of the manual activity level, but not influenced by the mental task load. High manual task loads resulted in a large percentage of mental errors even under low mental task loads. Although all the pilots gave similar subjective ratings when the manual task load was high, subjective ratings showed greater individual differences with high mental task loads. Altitude or airspeed deviations and subjective ratings were most correlated when the total task load was very high. Although airspeed deviations, altitude deviations, and subjective workload ratings were similar for both low experience and high experience pilots, at very high total task loads, mental performance was much lower for the low experience pilots.

Berg, S. L.↗

Psychophysiological Sensing and State Classification for Attention Management in Commercial Aviation

Attention-related human performance limiting states (AHPLS) can cause pilots to lose airplane state awareness (ASA), and their detection is important to improving commercial aviation safety. The Commercial Aviation Safety Team found that the majority of recent international commercial aviation accidents attributable to loss of control inflight involved flight crew loss of airplane state awareness, and that distraction of various forms was involved in all of them. Research on AHPLS, including channelized attention, diverted attention, startle / surprise, and confirmation bias, has been recommended in a Safety Enhancement (SE) entitled "Training for Attention Management." To accomplish the detection of such cognitive and psychophysiological states, a broad suite of sensors has been implemented to simultaneously measure their physiological markers during high fidelity flight simulation human subject studies. Pilot participants were asked to perform benchmark tasks and experimental flight scenarios designed to induce AHPLS. Pattern classification was employed to distinguish the AHPLS induced by the benchmark tasks. Unimodal classification using pre-processed electroencephalography (EEG) signals as input features to extreme gradient boosting, random forest and deep neural network multiclass classifiers was implemented. Multi-modal classification using galvanic skin response (GSR) in addition to the same EEG signals and using the same types of classifiers produced increased accuracy with respect to the unimodal case (90 percent vs. 86 percent), although only via the deep neural network classifier. These initial results are a first step toward the goal of demonstrating simultaneous real time classification of multiple states using multiple sensing modalities in high-fidelity flight simulators. This detection is intended to support and inform training methods under development to mitigate the loss of ASA and thus reduce accidents and incidents.

Harrivel, Angela R.↗

Measuring Pilot Workload in a Moving-base Simulator. Part 2: Building Levels of Workload

Pilot behavior in flight simulators often use a secondary task as an index of workload. His routine to regard flying as the primary task and some less complex task as the secondary task. While this assumption is quite reasonable for most secondary tasks used to study mental workload in aircraft, the treatment of flying a simulator through some carefully crafted flight scenario as a unitary task is less justified. The present research acknowledges that total mental workload depends upon the specific nature of the sub-tasks that a pilot must complete as a first approximation, flight tasks were divided into three levels of complexity. The simplest level (called the Base Level) requires elementary maneuvers that do not utilize all the degrees of freedom of which an aircraft, or a moving-base simulator; is capable. The second level (called the Paired Level) requires the pilot to simultaneously execute two Base Level tasks. The third level (called the Complex Level) imposes three simultaneous constraints upon the pilot.

Kantowitz, B. H.↗

A survey on checkpointing strategies: Should we always checkpoint à la Young/Daly?

The Young/Daly formula provides an approximation of the optimal checkpointing period for a parallel application executing on a supercomputing platform. It was originally designed to handle fail-stop errors for preemptible tightly-coupled applications, but has been extended to other application and resilience frameworks. Here, we provide some background and survey various scenarios to assess the usefulness and limitations of the formula, both for preemptible applications and workflow applications represented as a graph of tasks. We also discuss scenarios with uncertainties, and extend the study to silent errors. We exhibit cases where the optimal period is of a different order than that dictated by the Young/Daly formula, and finally we explain how checkpointing can be further combined with replication.

97 MATHEMATICS AND COMPUTING↗

HURON (HUman and Robotic Optimization Network) Multi-Agent Temporal Activity Planner/Scheduler

HURON solves the problem of how to optimize a plan and schedule for assigning multiple agents to a temporal sequence of actions (e.g., science tasks). Developed as a generic planning and scheduling tool, HURON has been used to optimize space mission surface operations. The tool has also been used to analyze lunar architectures for a variety of surface operational scenarios in order to maximize return on investment and productivity. These scenarios include numerous science activities performed by a diverse set of agents: humans, teleoperated rovers, and autonomous rovers. Once given a set of agents, activities, resources, resource constraints, temporal constraints, and de pendencies, HURON computes an optimal schedule that meets a specified goal (e.g., maximum productivity or minimum time), subject to the constraints. HURON performs planning and scheduling optimization as a graph search in state-space with forward progression. Each node in the graph contains a state instance. Starting with the initial node, a graph is automatically constructed with new successive nodes of each new state to explore. The optimization uses a set of pre-conditions and post-conditions to create the children states. The Python language was adopted to not only enable more agile development, but to also allow the domain experts to easily define their optimization models. A graphical user interface was also developed to facilitate real-time search information feedback and interaction by the operator in the search optimization process. The HURON package has many potential uses in the fields of Operations Research and Management Science where this technology applies to many commercial domains requiring optimization to reduce costs. For example, optimizing a fleet of transportation truck routes, aircraft flight scheduling, and other route-planning scenarios involving multiple agent task optimization would all benefit by using HURON.

Hua, Hook↗

NRAP Task 5: Preliminary Evaluation of the Cost of Responding to a Hypothetical Leakage Scenario Using the NRAP/SMART TALES Model and other NRAP Tools

Poster presentation illustrating the use of tools developed as part of Task 5 of Phase 3 of NRAP to estimate the technical performance and costs of implementing remedial responses to address a leak of fluid out of the storage formation at a CO2 saline storage project. Presented at the 2024 FECM - NETL Carbon Management Research Project Review Meeting, 5-9 August, 2024, Pittsburgh, PA.

Morgan, David↗

Configuration and Projected Capabilities of the Common Habitat Medical Care Facility

The Common Habitat is a large, long-duration habitat being explored as part of a conceptual study (not an active NASA program) that uses an SLS core stage Liquid Oxygen (LOX) tank as its primary structure. It is intended for use on the Moon as part of a permanently occupied outpost, on Mars as part of an outpost that will be occupied for hundreds of days at a time, and in deep space as part of the Deep Space Exploration Vehicle where it will support crewed missions up to 1200 days in duration. A study of internal orientation and crew size resulted in a Common Habitat configuration sized for a crew of eight with a three-deck horizontal orientation. Additional work outside the scope of this paper is developing a vertical translation system, a crew mobility aids system based on wearable gecko-derived grippers, and a crew seating/restraint system. These systems are all assumed for use in conjunction with the Medical Care Facility, which is needed to maintain crew well-being during these missions, where distance from Earth precludes the possibility of evacuation to Earth. This paper describes recent improvements in the Common Habitat Medical Care Facility and associated benefits for crew survivability in long duration missions beyond Earth orbit. These improvements were made with the assistance of a NASA Pathways intern whose experience includes a tour of duty in Afghanistan as an Army combat medic with the 691st GHOST-T, attached to the 1st and 7th US Special Forces Groups as part of Operation Freedom’s Sentinel, where he helped provide far-forward surgical capabilities in austere combat environments. The initial baseline Medical Care Facility was developed working in conjunction with University of Houston Space Architecture graduate students. The facility was placed on the upper deck of the Common Habitat in a location that provided privacy, operational volume, and was close to the vertical translation pathway. The notional outfitting repurposed component CAD models from unrelated studies and notionally indicated a level of care roughly equivalent to that aboard the International Space Station. The CAD modeling provided notional stowage volumes, a deployable surface, some fixed equipment, an ultrasound, and a potentially reconfigurable treatment table. While this facility is clearly a competent arrangement, it was desired to leverage available expertise and upgrade the station given the vast distances from Earth to be experienced by the Common Habitat. Key driving requirements applied to the upgrade included to provide Medical Level of Care V, offer enhanced telemedicine capabilities, provide patient physical accommodation, provide caregiver access to the patient from all sides, include sliding pocket doors for access to hygiene and to the Vertical Translation System, and to add any additional capability possible for the best achievable medical care. The first step in the facility upgrade was to quantify the current medical inventory on the International Space Station and ensure that sufficient stowage volume was present for this purpose. To that end, the ISS medical kits were reviewed, and eight full size mid deck lockers were placed in the facility. A number of additional devices were also added, based on the intern’s combat medic experience. Also, two fixed shelves and one horizontal work surface were added to the Medical Care Facility, with the shelves providing storage space for the additional devices and the work surface providing a location for the caregiver to work or stage equipment. Four display monitors were added to the wall above the horizontal work surface, supporting data display, telemedicine, conferencing, or other needs. The existing treatment table was replaced with a mobile surgical stretcher-chair. Two additional doors were added to the Medical Care Facility. One leads directly to the hygiene compartment, allowing it to support medical operations in addition to providing galley/wardroom support. The other door leads directly into the Vertical Translation System. The wall adjacent to the subsystems bay was moved, adding additional volume to the Medical Care Facility. This improved caregiver access to the patient and allowed for a larger number of caregivers to be present. It also provided options for relocation of support equipment relative to the patient as needed. In the upgraded Medical Care Facility, the Surgical Stretcher-Chair and the Vertical Translation System can work together to provide incapacitated crew member transport from a site of injury on any deck of the Common Habitat to the Medical Care Facility. It can also support patient treatment in a variety of positions including a variety of sitting postures and a supine posture at a variety of pitch angles. The facility can also support caregiver office work for review of examination results, private consultation, inventory and maintenance, and a variety of other purposes. A forward activity will be to conduct evaluations of the Medical Care Facility with different medical scenarios. Additionally, ambient and task lighting selections remain as forward work. The eight mid deck lockers can be augmented to use as portable equipment carts, similar to a manner in which maintenance facility stowage was used as portable carts during the NASA Desert Research and Technology Studies in the Constellation Program. Trash accommodation will also need forward work to assess, including provision for wet trash, dry trash, and biological waste. It will be important to assess a redesign of the surgical stretcher-chair. The commercial version used in the upgrade can only enable vertical translation in the seated configuration, requiring the patient to bend both hips and knees. A possible redesign of the chair will allow for vertical translation without requiring any bending at the hip or knees. Also, the commercial version is wheeled, making it mobile in gravity but unanchored in microgravity. Work will be needed to adapt the chair for gravity-independent performance. The hygiene compartment can be redesigned for dual-use medical scrub and galley handwash facility. Pending sufficient volume, it may also be possible to place sanitation equipment in this location to clean medical tools. Finally, most space architectures have never allowed for more than one incapacitated crew member, but several scenarios could potentially injure two or more crew in the same incident. This facility could be assessed to determine its present ability to address two or more injured crew in parallel and determine the potential upper limit for number of treatable crew in a multi-crew injury scenario, or to treat polytrauma of a single patient.

Habitat↗

Configuration and Projected Capabilities of the Common Habitat Medical Care Facility

The Common Habitat is a large, long-duration habitat being explored as part of a conceptual study (not an active NASA program) that uses an SLS core stage Liquid Oxygen (LOX) tank as its primary structure. It is intended for use on the Moon as part of a permanently occupied outpost, on Mars as part of an outpost that will be occupied for hundreds of days at a time, and in deep space as part of the Deep Space Exploration Vehicle where it will support crewed missions up to 1200 days in duration. A study of internal orientation and crew size resulted in a Common Habitat configuration sized for a crew of eight with a three-deck horizontal orientation. Additional work outside the scope of this paper is developing a vertical translation system, a crew mobility aids system based on wearable gecko-derived grippers, and a crew seating/restraint system. These systems are all assumed for use in conjunction with the Medical Care Facility, which is needed to maintain crew well-being during these missions, where distance from Earth precludes the possibility of evacuation to Earth. This paper describes recent improvements in the Common Habitat Medical Care Facility and associated benefits for crew survivability in long duration missions beyond Earth orbit. These improvements were made with the assistance of a NASA Pathways intern whose experience includes a tour of duty in Afghanistan as an Army combat medic with the 691st GHOST-T, attached to the 1st and 7th US Special Forces Groups as part of Operation Freedom’s Sentinel, where he helped provide far-forward surgical capabilities in austere combat environments. The initial baseline Medical Care Facility was developed working in conjunction with University of Houston Space Architecture graduate students. The facility was placed on the upper deck of the Common Habitat in a location that provided privacy, operational volume, and was close to the vertical translation pathway. The notional outfitting repurposed component CAD models from unrelated studies and notionally indicated a level of care roughly equivalent to that aboard the International Space Station. The CAD modeling provided notional stowage volumes, a deployable surface, some fixed equipment, an ultrasound, and a potentially reconfigurable treatment table. While this facility is clearly a competent arrangement, it was desired to leverage available expertise and upgrade the station given the vast distances from Earth to be experienced by the Common Habitat. Key driving requirements applied to the upgrade included to provide NASA’s Medical Level of Care V, offer enhanced telemedicine capabilities, provide patient physical accommodation, provide caregiver access to the patient from all sides, include sliding pocket doors for access to hygiene and to the Vertical Translation System, and to add any additional capability possible for the best achievable medical care. The first step in the facility upgrade was to quantify the current medical inventory on the International Space Station and ensure that sufficient stowage volume was present for this purpose. To that end, the ISS medical kits were reviewed, and eight full size mid deck lockers were placed in the facility. A number of additional devices were also added, based on the co-author’s combat medic experience. Also, two fixed shelves and one horizontal work surface were added to the Medical Care Facility, with the shelves providing storage space for the additional devices and the work surface providing a location for the caregiver to work or stage equipment. Four display monitors were added to the wall above the horizontal work surface, supporting data display, telemedicine, conferencing, or other needs. The existing treatment table was replaced with a mobile surgical stretcher-chair. Two additional doors were added to the Medical Care Facility. One leads directly to the hygiene compartment, allowing it to support medical operations in addition to providing galley/wardroom support. The other door leads directly into the Vertical Translation System. The wall adjacent to the subsystems bay was moved, adding additional volume to the Medical Care Facility. This improved caregiver access to the patient and allowed for a larger number of caregivers to be present. It also provided options for relocation of support equipment relative to the patient as needed. In the upgraded Medical Care Facility, the Surgical Stretcher-Chair and the Vertical Translation System can work together to provide incapacitated crew member transport from a site of injury on any deck of the Common Habitat to the Medical Care Facility. It can also support patient treatment in a variety of positions including a variety of sitting postures and a supine posture at a variety of pitch angles. The facility can also support caregiver office work for review of examination results, private consultation, inventory and maintenance, and a variety of other purposes. A forward activity will be to conduct evaluations of the Medical Care Facility with different medical scenarios. Additionally, ambient and task lighting selections remain as forward work. The eight middeck lockers can be augmented to use as portable equipment carts, similar to a manner in which maintenance facility stowage was used as portable carts during the NASA Desert Research and Technology Studies in the Constellation Program. Trash accommodation will also need forward work to assess, including provision for wet trash, dry trash, and biological waste. It will be important to assess a redesign of the surgical stretcher-chair. The commercial version used in the upgrade can only enable vertical translation in the seated configuration, requiring the patient to bend both hips and knees. A possible redesign of the chair will allow for vertical translation without requiring any bending at the hip or knees. Also, the commercial version is wheeled, making it mobile in gravity but unanchored in microgravity. Work will be needed to adapt the chair for gravity-independent performance. The hygiene compartment can be redesigned to serve both as a medical scrub facility and for galley hand washing. Pending sufficient volume, it may also be possible to place sanitation equipment in this location to clean medical tools. Finally, most space architectures have never allowed for more than one incapacitated crew member, but several scenarios could potentially injure two or more crew in the same incident. This facility could be assessed to determine its present ability to address two or more injured crew in parallel and determine the potential upper limit for number of treatable crew in a multi-crew injury scenario, or to treat polytrauma of a single patient.

Common Habitat↗

Reactor System Demonstration with Cyber-Attack Scenarios Using CrowPis and Arduino Microcontrollers

This study covers developing and simulating nuclear reactor system using CrowPis and Arduino microcontrollers for demonstrating cyber-attack scenarios. The team was tasked with implementing more sensors and cybersecurity aspects to the reactor program that was created by last year’s high school interns. The team received the opportunity to collaborate and obtain advice from multiple university interns that helped us gain a better perspective of our project. Our mentor’s background in nuclear science was pivotal to our understanding what we could add to the reactor program to make it as realistic as possible. The first week of our internship was spent reading as much material as possible to gain an understanding of and the background for cyber-attacks and nuclear science. Nuclear science was a new horizon for each of the high school interns on the team, so spending this time in the beginning of our internship was crucial to our success. For the remaining portion of our internship, the team collectively did our best to implement as many sensors and use as much hardware as we could to make an accurate representation of a nuclear reactor in the program that was created. This internship was a big learning experience for everyone on the team. We all gained so many insights into the nuclear world and how it can benefit our lives, as well as how so many moving pieces are needed for it to work properly.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗