Search NASA⌕ Search

SEARCH · Search NASA

Results for “Human-in-the-Loop Testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

NASA Processes and Requirements for Conducting Human-in-the-Loop Closed Chamber Tests

NASA has specific processes and requirements that must be followed for tests involving human subjects to be conducted in a safe and effective manner. There are five distinct phases of test operations. Phase one, the test request phase, consists of those activities related to initiating, processing, reviewing, and evaluating the test request. Phase two, the test preparation phase consists of those activities related to planning, coordinating, documenting, and building up the test. Phase three, the test readiness phase consists of those activities related to verifying and reviewing the planned test operations. Phase four, the test activity phase, consists of all pretest operations, functional checkouts, emergency drills, and test operations. Phase five, the post test activity phase, consists of those activities performed once the test is completed, including briefings, documentation of anomalies, data reduction and archiving, and reporting. Project management processes must be followed for facility modifications and major test buildup, which include six phases: initiation and assessment, requirements evaluation, preliminary design, detailed design, use readiness review (URR) and acceptance. Compliance with requirements for safety and quality assurance are documented throughout the test buildup and test operation processes. Tests involving human subjects must be reviewed by the applicable Institutional Review Board (IRB).

Barta, Daniel J.↗

Formal Verification of a Conflict Resolution and Recovery Algorithm

New air traffic management concepts distribute the duty of traffic separation among system participants. As a consequence, these concepts have a greater dependency and rely heavily on on-board software and hardware systems. One example of a new on-board capability in a distributed air traffic management system is air traffic conflict detection and resolution (CD&R). Traditional methods for safety assessment such as human-in-the-loop simulations, testing, and flight experiments may not be sufficient for this highly distributed system as the set of possible scenarios is too large to have a reasonable coverage. This paper proposes a new method for the safety assessment of avionics systems that makes use of formal methods to drive the development of critical systems. As a case study of this approach, the mechanical veri.cation of an algorithm for air traffic conflict resolution and recovery called RR3D is presented. The RR3D algorithm uses a geometric optimization technique to provide a choice of resolution and recovery maneuvers. If the aircraft adheres to these maneuvers, they will bring the aircraft out of conflict and the aircraft will follow a conflict-free path to its original destination. Veri.cation of RR3D is carried out using the Prototype Verification System (PVS).

Maddalon, Jeffrey↗

Overview of Past and Future EVA Suit Testing and Training in Chamber B at NASA JSC

The Crew and Thermal System Division’s Chamber B has received an increase in demand for next generation suit testing. Chamber B is the National Aeronautics and Space Administration (NASA) Johnson Space Center’s only human-rated thermal vacuum (TVAC) chamber. Historically it was used in Gemini, Apollo, Skylab, Shuttle, and the International Space Station (ISS) suit tests. Recently, the chamber has been returned to service with new capabilities for suit testing. In 2023, The Exploration Extravehicular Mobility Unit (xEMU) underwent a 5-day, extensive thermal vacuum test that included both a full-bodied Exploration Pressure Garment Suit (xPGS) as well as a high fidelity Short xEMU. SpaceX has also used the chamber for human-in-the-loop (HITL) qualification and acceptance testing on their flight Polaris Dawn suits for the first-ever commercial extravehicular activity (EVA). Current chamber Manlock B2 upgrades include the support of two test subjects at the same time with updated chamber interface and support systems for the next generation of suits. This paper will discuss the history of spacesuit testing in the chamber, the recent testing for commercial and NASA suits, and upgrades to accommodate new test requirements.

Human-in-the-loop↗

Overview of Past and Future EVA Suit Testing and Training in Chamber B at NASA JSC

The Crew and Thermal System Division’s Chamber B has received an increase in demand for next generation suit testing. Chamber B is the National Aeronautics and Space Administration (NASA) Johnson Space Center’s only human-rated thermal vacuum (TVAC) chamber. Historically it was used in Gemini, Apollo, Skylab, Shuttle, and the International Space Station (ISS) suit tests. Recently, the chamber has been returned to service with new capabilities for suit testing. In 2023, The Exploration Extravehicular Mobility Unit (xEMU) underwent a 5-day, extensive thermal vacuum test that included both a full-bodied Exploration Pressure Garment Suit (xPGS) as well as a high fidelity Short xEMU. SpaceX has also used the chamber for human-in-the-loop (HITL) qualification and acceptance testing on their flight Polaris Dawn suits for the first-ever commercial extravehicular activity (EVA). Current chamber Manlock B2 upgrades include the support of two test subjects at the same time with updated chamber interface and support systems for the next generation of suits. This paper will discuss the history of spacesuit testing in the chamber, the recent testing for commercial and NASA suits, and upgrades to accommodate new test requirements.

Polaris Dawn↗

Design of a Multi-mode Flight Deck Decision Support System for Airborne Conflict Management

NASA Langley has developed a multi-mode decision support system for pilots operating in a Distributed Air-Ground Traffic Management (DAG-TM) environment. An Autonomous Operations Planner (AOP) assists pilots in performing separation assurance functions, including conflict detection, prevention, and resolution. Ongoing AOP design has been based on a comprehensive human factors analysis and evaluation results from previous human-in-the-loop experiments with airline pilot test subjects. AOP considers complex flight mode interactions and provides flight guidance to pilots consistent with the current aircraft control state. Pilots communicate goals to AOP by setting system preferences and actively probing potential trajectories for conflicts. To minimize training requirements and improve operational use, AOP design leverages existing alerting philosophies, displays, and crew interfaces common on commercial aircraft. Future work will consider trajectory prediction uncertainties, integration with the TCAS collision avoidance system, and will incorporate enhancements based on an upcoming air-ground coordination experiment.

Barhydt, Richard↗

Conflict Resolution Performance in an Experimental Study of En Route Free Maneuvering Operations

NASA has developed a far-term air traffic management concept, termed Distributed Air/Ground Traffic Management (DAG-TM). One component of DAG-TM, En Route Free Maneuvering, allows properly trained flight crews of equipped autonomous aircraft to assume responsibility for separation from other autonomous aircraft and from Instrument Flight Rules (IFR) aircraft. Ground-based air traffic controllers continue to separate IFR traffic and issue flow management constraints to all aircraft. To examine En Route Free Maneuvering operations, a joint human-in-the-loop experiment was conducted in summer 2004 at the NASA Ames and Langley Research Centers. Test subject pilots used desktop flight simulators to resolve traffic conflicts and adhere to air traffic flow constraints issued by subject controllers. The experimental airspace integrated both autonomous and IFR aircraft at varying traffic densities. This paper presents a subset of the En Route Free Maneuvering experimental results, focusing on airborne and ground-based conflict resolution, and the effects of increased traffic levels on the ability of pilots and air traffic controllers to perform this task. The results show that, in general, increases in autonomous traffic do not significantly impact conflict resolution performance. In addition, pilot acceptability of autonomous operations remains high throughout the range of traffic densities studied. Together with previously reported findings, these results continue to support the feasibility of the En Route Free Maneuvering component of DAG-TM.

Doble, Nathan A.↗

Airborne Management of Traffic Conflicts in Descent With Arrival Constraints

NASA is studying far-term air traffic management concepts that may increase operational efficiency through a redistribution of decisionmaking authority among airborne and ground-based elements of the air transportation system. One component of this research, En Route Free Maneuvering, allows trained pilots of equipped autonomous aircraft to assume responsibility for traffic separation. Ground-based air traffic controllers would continue to separate traffic unequipped for autonomous operations and would issue flow management constraints to all aircraft. To evaluate En Route Free Maneuvering operations, a human-in-the-loop experiment was jointly conducted by the NASA Ames and Langley Research Centers. In this experiment, test subject pilots used desktop flight simulators to resolve conflicts in cruise and descent, and to adhere to air traffic flow constraints issued by test subject controllers. Simulators at NASA Langley were equipped with a prototype Autonomous Operations Planner (AOP) flight deck toolset to assist pilots with conflict management and constraint compliance tasks. Results from the experiment are presented, focusing specifically on operations during the initial descent into the terminal area. Airborne conflict resolution performance in descent, conformance to traffic flow management constraints, and the effects of conflicting traffic on constraint conformance are all presented. Subjective data from subject pilots are also presented, showing perceived levels of workload, safety, and acceptability of autonomous arrival operations. Finally, potential AOP functionality enhancements are discussed along with suggestions to improve arrival procedures.

Doble, Nathan A.↗

Usability Study of Two Collocated Prototype System Displays

Currently, most of the displays in control rooms can be categorized as status screens, alerts/procedures screens (or paper), or control screens (where the state of a component is changed by the operator). The primary focus of this line of research is to determine which pieces of information (status, alerts/procedures, and control) should be collocated. Two collocated displays were tested for ease of understanding in an automated desktop survey. This usability study was conducted as a prelude to a larger human-in-the-loop experiment in order to verify that the 2 new collocated displays were easy to learn and usable. The results indicate that while the DC display was preferred and yielded better performance than the MDO display, both collocated displays can be easily learned and used.

Trujillo, Anna C.↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Development of a Free-Flight Simulation Infrastructure

In anticipation of a projected rise in demand for air transportation, NASA and the FAA are researching new air-traffic-management (ATM) concepts that fall under the paradigm known broadly as ":free flight". This paper documents the software development and engineering efforts in progress by Seagull Technology, to develop a free-flight simulation (FFSIM) that is intended to help NASA researchers test mature-state concepts for free flight, otherwise referred to in this paper as distributed air / ground traffic management (DAG TM). Under development is a distributed, human-in-the-loop simulation tool that is comprehensive in its consideration of current and envisioned communication, navigation and surveillance (CNS) components, and will allow evaluation of critical air and ground traffic management technologies from an overall systems perspective. The FFSIM infrastructure is designed to incorporate all three major components of the ATM triad: aircraft flight decks, air traffic control (ATC), and (eventually) airline operational control (AOC) centers.

Miles, Eric S.↗

Data Management for Mars Exploration Rovers

Data Management for the Mars Exploration Rovers (MER) project is a comprehensive system addressing the needs of development, test, and operations phases of the mission. During development of flight software, including the science software, the data management system can be simulated using any POSIX file system. During testing, the on-board file system can be bit compared with files on the ground to verify proper behavior and end-to-end data flows. During mission operations, end-to-end accountability of data products is supported, from science observation concept to data products within the permanent ground repository. Automated and human-in-the-loop ground tools allow decisions regarding retransmitting, re-prioritizing, and deleting data products to be made using higher level information than is available to a protocol-stack approach such as the CCSDS File Delivery Protocol (CFDP).

Mars missions↗

Pilot Preference, Compliance, and Performance With an Airborne Conflict Management Toolset

A human-in-the-loop experiment was conducted at the NASA Ames and Langley Research Centers, investigating the En Route Free Maneuvering component of a future air traffic management concept termed Distributed Air/Ground Traffic Management (DAG-TM). NASA Langley test subject pilots used the Autonomous Operations Planner (AOP) airborne toolset to detect and resolve traffic conflicts, interacting with subject pilots and air traffic controllers at NASA Ames. Experimental results are presented, focusing on conflict resolution maneuver choices, AOP resolution guidance acceptability, and performance metrics. Based on these results, suggestions are made to further improve the AOP interface and functionality.

Doble, Nathan A.↗

Speeding-up fuzzing through directional seeds

Abstract Fuzzing is an automated process for discovering inputs in a program that may trigger unexpected behavior. Today, fuzzing has become a standard practice for the discovery of bugs and security vulnerabilities. However, the main issue with such practices is that the exploration of the input space of programs can often be prohibitively expensive. Therefore, several alternative fuzzing strategies have been introduced during the last few years. Some fuzzing techniques rely on human expertise to provide a plausible set of initial input examples, namely, seeds. However, the process of handcrafting seeds for fuzzing purposes often becomes strenuous for humans as it requires a deeper understanding of the Program-Under-Test (PUT). Also, the use of known inputs to programs often does not trigger vulnerable program behavior or may not reach potentially vulnerable code locations. To address those issues, we propose a seed generation framework that enables Human-In-The-Loop (HITL) directed fuzzing where the human assumes a more active role in the creation of seeds that can penetrate and assess desired locations of the PUT. Our proposed framework uses Symbolic Execution (SE) to generate seeds that exercise paths to target program locations. Moreover, our framework enables the visualization of the explored execution paths in the binary of the PUT for the generated seeds. We evaluated our approach on a set of 12 carefully designed C programs with diverse characteristics that mimic real-world programs. The experimental results show the effectiveness of the proposed approach in improving the performance of standard fuzzing tools such as the American Fuzzy Lop ("Image missing" <#comment/> ). Specifically, our solution can generate seeds that substantially enhance the performance of the fuzzer, achieving speedups ranging from $$1.46\times $$ 1.46 × to $$68.53\times $$ 68.53 × for branch conditions, $$1.39\times $$ 1.39 × to $$254.62\times $$ 254.62 × for branch depths, $$14,879.59\times $$ 14 , 879.59 × to $$30,295.88\times $$ 30 , 295.88 × for branch widths over traditional seeds. Additionally, the speedup increases with the number of target function ranging from $$12,260\times $$ 12 , 260 × to $$22,856.07\times $$ 22 , 856.07 × over traditional seeds while only requiring less than 15 seconds on average for the seed generation step.

97 MATHEMATICS AND COMPUTING↗

Agentic traffic intelligence: Augmented human-in-the-loop scenario generation for microscopic traffic simulation

Traditional microscopic traffic simulation generation often relies on static datasets and manual design, limiting its ability to simulate complex conditions easily. This paper presents a novel framework, Agentic Traffic Intelligence, which combines human approval large language models (LLMs), the Real-Twin tool, and multi-agent systems to perform realistic microscopic traffic simulation scenario generation. The proposed framework incorporates human-in-the-loop (HIL) control, retrieval-augmented generation (RAG), and multi-agent control mechanisms. HIL mechanisms are used to guide multiple LLMs focused on attributes for microscopic simulation generation and to improve the interpretability and transparency of LLM execution for users. RAG enhances context extraction by dynamically integrating external knowledge sources for traffic scenario generation foundations. A multi-agent architecture with supervisory control coordinates the interaction of simulation components, including traffic simulators, control logic, and calibration tools. This enables the synthesis of simulation-ready scenarios that reflect dynamic demand profiles and behavior controls. Furthermore, the framework fuses multisource traffic data with unstructured context and supports iterative refinement through interactive user feedback. Validated through microscopic simulation using Simulation of Urban Mobility, the generated scenarios demonstrate high-fidelity network generation with inflow and turn movement and behavioral calibration, offering a robust and efficient tool for stress-testing and optimizing urban mobility systems.

Hierarchical multi-agent control↗

Evaluation of Airborne Precision Spacing in a Human-in-the-Loop Experiment

A significant bottleneck in the current air traffic system occurs at the runway. Expanding airports and adding new runways will help solve this problem; however, this comes with significant costs: financially, politically and environmentally. A complementary solution is to safely increase the capacity of current runways. This can be achieved by precisely spacing aircraft at the runway threshold, with a resulting reduction in the spacing bu er required under today s operations. At NASA's Langley Research Center, the Airspace Systems program has been investigating airborne technologies and procedures that will assist the flight crew in achieving precise spacing behind another aircraft. A new spacing clearance allows the pilot to follow speed cues from a new on-board guidance system called Airborne Merging and Spacing for Terminal Arrivals (AMSTAR). AMSTAR receives Automatic Dependent Surveillance-Broadcast (ADS-B) reports from an assigned, leading aircraft and calculates the appropriate speed for the ownship to fly to achieve the desired spacing interval, time- or distance-based, at the runway threshold. Since the goal is overall system capacity, the speed guidance algorithm is designed to provide system-wide benefits and stability to a string of arriving aircraft. An experiment was recently performed at the NASA Langley Air Traffic Operations Laboratory (ATOL) to test the flexibility of Airborne Precision Spacing operations under a variety of operational conditions. These included several types of merge and approach geometries along with the complementary merging and in-trail operations. Twelve airline pilots and four controllers participated in this simulation. Performance and questionnaire data were collected from a total of eighty-four individual arrivals. The pilots were able to achieve precise spacing with a mean error of 0.5 seconds and a standard deviation of 4.7 seconds. No statistically significant di erences in spacing performance were found between in-trail and merging operations or among the three modeled airspaces. Questionnaire data showed general acceptance for both pilots and controllers. These results reinforce previous findings from full-mission simulation and flight evaluation of the in-trail operations. This paper reviews the results of this simulation in detail.

Barmore, Bryan E.↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Autonomous Flow Electrochemistry for Accelerated Catalyst Discovery

Our objective is to develop an Autonomous Chemical Experimentation (ACE) platform that accelerates discovery of new catalytic transformations and other energy-relevant chemical reactions and processes. We intentionally designed ACE to be highly modular, both with respect to its rapid deployment to different chemistries and experimental workflows as well as incorporation of a wide range of different AI algorithms. In addition to the development of the core software architecture, initial efforts were made to incorporate Large Language Models to provide human-interpretable reasoning of the optimizer’s actions, and to develop a user-friendly graphical interface for experimental researchers. ACE was demonstrated using a flow electrocatalysis platform containing an inline FTIR spectrometer for real-time analysis and quantification of the reaction outcome. Human-in-the-loop experiments were performed in which a human researcher conducted an experiment using electrode potentials suggested by ACE, then fed the spectral data back to ACE for decision making. After confirming the successful function of the optimizer, efforts were next directed to automation of the hardware and performed full autonomy tests using three reactions: catalytic oxidation of formate, catalytic oxidation of cyclohexanol, and oxidation of hydroquinone. These studies confirm that ACE can close the loop between reaction execution, analysis, and optimization. They also reveal that more improved product detection methods will be essential for ACE to make well-informed decisions for reactions with low conversions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A knowledge-based approach to identification and adaptation in dynamical systems control

Artificial intelligence techniques are applied to the problems of model form and parameter identification of large-scale dynamic systems. The object-oriented knowledge representation is discussed in the context of causal modeling and qualitative reasoning. Structured sets of rules are used for implementing qualitative component simulations, for catching qualitative discrepancies and quantitative bound violations, and for making reconfiguration and control decisions that affect the physical system. These decisions are executed by backward-chaining through a knowledge base of control action tasks. This approach was implemented for two examples: a triple quadrupole mass spectrometer and a two-phase thermal testbed. Results of tests with both of these systems demonstrate that the software replicates some or most of the functionality of a human operator, thereby reducing the need for a human-in-the-loop in the lower levels of control of these complex systems.

Glass, B. J.↗