Search NASA⌕ Search

SEARCH · Search NASA

Results for “Human-in-the-Loop Testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Human and Technology Integration Evaluation of Advanced Automation and Data Visualization

While the existing United States (U.S.) light water reactors are highly reliable, safe, and provide a significant proportion of carbon-free electricity, the cost of operating and maintaining them has become less competitive compared to other electricity generating sources. The reason for the gap in operating and maintenance (O&M) costs can be at least in part attributed to the advent of new digital technologies that other electricity generating industries are currently using. Advanced capabilities including digital instrumentation and control (I&C) systems, advanced automation and analytics, and greater span of data integration (i.e., connectedness) across these non-nuclear plants has transformed the way work is performed and ultimately given them a competitive advantage in terms of the cost required for operating, maintaining, and supporting them. To reduce O&M cost and address obsolescence of the aging I&C infrastructure of the existing U.S. light water reactors, the U.S. Department of Energy (DOE) Light Water Reactor Sustainability (LWRS) Program Plant Modernization Pathway is conducting targeting multidisciplinary research that 1) delivers a sustainable business model to enable a cost-competitive U.S. nuclear industry and 2) is developing technology modernization solutions that address aging and obsolescence challenges. The work described in this report supports these two objectives and describes the demonstration of human and technology integration across recent industry collaborations to support their large-scale digital I&C modifications. This technical report describes the demonstration of the human and technology integration methodology in performing full-scale performance-based human-in-the-loop tests to evaluate plant-specific advanced automation and data visualization applications within these collaborators’ digital modifications. This technical report also documents future applications of human and technology integration that expand beyond main control room modernization and digital I&C upgrades, which have been a central focus to date. Thus, this technical report discusses how to implement human and technology integration across new business opportunities and how to develop an evaluation plan that defines measures and criteria, and documents key assumptions to support full plant modernization.

99 GENERAL AND MISCELLANEOUS↗

Human Dimensions of Energy Systems Workshop - September 6-7, 2022: Findings and Next Steps

Advanced energy technologies and informed policies are necessary but not sufficient to accelerate the energy transition at the speed required to meet our shared climate and energy resilience goals. To fully understand and implement integrated energy systems, we need to be able to model and analyze the complete system-of-systems, including the behavior of people interacting with and being impacted by energy systems. The purpose of this workshop is: Understand how human behavior and decisions affect the performance of energy systems with a focus on resilience and human well-being; Identify opportunities to improve energy system design to explicitly consider human behavior and well-being; Enhance our capability to model and predict human actions. Develop the ability to stress test these integrated systems; Identify or develop requirements and tools for more robust system designs that can be used for human-in-the-loop exercises and training; Identify opportunities for joint research, joint appointments, and collaboration; and Guide internal investments and strategic hires.

analyze↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Speeding-up fuzzing through directional seeds

Abstract Fuzzing is an automated process for discovering inputs in a program that may trigger unexpected behavior. Today, fuzzing has become a standard practice for the discovery of bugs and security vulnerabilities. However, the main issue with such practices is that the exploration of the input space of programs can often be prohibitively expensive. Therefore, several alternative fuzzing strategies have been introduced during the last few years. Some fuzzing techniques rely on human expertise to provide a plausible set of initial input examples, namely, seeds. However, the process of handcrafting seeds for fuzzing purposes often becomes strenuous for humans as it requires a deeper understanding of the Program-Under-Test (PUT). Also, the use of known inputs to programs often does not trigger vulnerable program behavior or may not reach potentially vulnerable code locations. To address those issues, we propose a seed generation framework that enables Human-In-The-Loop (HITL) directed fuzzing where the human assumes a more active role in the creation of seeds that can penetrate and assess desired locations of the PUT. Our proposed framework uses Symbolic Execution (SE) to generate seeds that exercise paths to target program locations. Moreover, our framework enables the visualization of the explored execution paths in the binary of the PUT for the generated seeds. We evaluated our approach on a set of 12 carefully designed C programs with diverse characteristics that mimic real-world programs. The experimental results show the effectiveness of the proposed approach in improving the performance of standard fuzzing tools such as the American Fuzzy Lop ("Image missing" <#comment/> ). Specifically, our solution can generate seeds that substantially enhance the performance of the fuzzer, achieving speedups ranging from $$1.46\times $$ 1.46 × to $$68.53\times $$ 68.53 × for branch conditions, $$1.39\times $$ 1.39 × to $$254.62\times $$ 254.62 × for branch depths, $$14,879.59\times $$ 14 , 879.59 × to $$30,295.88\times $$ 30 , 295.88 × for branch widths over traditional seeds. Additionally, the speedup increases with the number of target function ranging from $$12,260\times $$ 12 , 260 × to $$22,856.07\times $$ 22 , 856.07 × over traditional seeds while only requiring less than 15 seconds on average for the seed generation step.

97 MATHEMATICS AND COMPUTING↗

Human-in-the-Loop Motion Control of a Two-DOF Hydraulic Backhoe Powered by the Hybrid Hydraulic Electric Architecture (HHEA)

Abstract The Hybrid Hydraulic-Electric Architecture (HHEA) combines the respective power density and control advantages of hydraulic and electric actuation to save energy for off-road vehicles. It uses a set of selectable common pressure rails to transmit the majority of power and electric actuation to modulate that power. As it is critical that off-road vehicles can perform tasks dexterously and exactly as commanded by the operator, the switchings between discrete pressure rails pose a potential challenge for smooth and precise motion. A control strategy consisting of a backstepping nominal control and least norm transition control has previously been developed to address this issue. It has been tested on 1 degree of freedom (DOF) testbeds where known trajectories were able to be tracked precisely. This paper presents the implementation of the HHEA motion control strategy on a 2-DOF backhoe operated by a human operator via a 2-DOF joystick. Unlike previous studies, the duty cycle is unknown beforehand and the decision to change pressure rails is taken in real-time. The efficacy of the motion control strategy has been validated experimentally. Several strategies to improve the user interface: control in workspace coordinates, pressure feedback, and velocity field-based task specification, have also been implemented and demonstrated to make operating the multiple DOF, HHEA actuated machine more intuitive to novice operators.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explainable Artificial Intelligence Technology for Predictive Maintenance

The domestic nuclear power plant fleet has relied on labor-intensive and time-consuming preventive maintenance programs, thus driving up operation and maintenance costs to achieve high-capacity factors. Artificial intelligence and machine learning can help simplify complex problems, such as diagnosing equipment degradation, to enable more effective decision-making. Benefits will be felt not only within existing analog and digital instrumentation and control, but also work processes, the integration of people with technology, and most importantly, the business case. Together, these hold promise to make nuclear power more efficient and reduce costs associated with operation and maintenance. While the artificial intelligence and machine learning technologies hold significant promise in the nuclear industry, there are challenges or barriers to their adoption. This report outlines the those different machine learning adoption barriers (categorized as historical, technical, economic, regulatory, and user) that the industry must overcome to realize the full benefits of artificial intelligence and machine learning capabilities for long-term economic sustainability. This report also provides solutions for some of these barriers by focusing on improving the explainability of machine learning to encourage trust from the end-user. Trust and explainability are essential to machine learning adoption. This report focuses on research-developed solutions to some of these barriers while analyzing a non-safety-related system, namely the circulating water system. This system frequently experiences waterbox fouling which our models preemptively diagnoses then explains to the operator how those conclusions were reached. This report presents and discusses the inherent trade-off between machine learning performance (in terms of accuracy) and explainability, where highly accurate machine learning methods (such as deep-learning) are the least explainable, and the most explainable methods (such as decision trees) are the least accurate. In addition, explainability of artificial intelligence techniques in terms of transparency and post-hoc metrics are discussed. This report outlines the importance of data novelty and value of new information in evaluating both the explainability and trustworthiness. Novelty detection helps to establish consistency or inconsistency of the new data with respect to the training data. On the other hand, value of information could be a part of the user-centric visualization recommendation system that request additional information to be collected, thereby strengthening the machine learning outcomes. During this project, a copyrighted user-centric visualization that aligns with a human-in-the-loop approach was developed. The user-centric visualization presents different levels of information and can be tailored as per user credentials to gain user confidence. One of the salient features of the user-centric visualization is it presents machine learning methods with explainability metrics. A simplified version of the user-centric visualization was presented to 32 users with varying levels of machine learning expertise. Feedback was solicited to test the hypothesis that the app contained sufficient explainability and that the users would trust the algorithm. Overall, the app was positively received, and the hypothesis was supported. This report discusses the trust-but-verify framework – a potential approach to build user trust artificial intelligence. The framework discusses trust from the human level to artificial intelligence level. The fundamental premise of the trust but verify framework is derived from an observation of nuclear safety culture (i.e., nuclear power plant personnel do not rely on a singular source of data to make a decision). This also ties back to the user-centric visualization that presents different levels of information to achieve both explainability and trustworthiness of artificial intelligence. Even so, the adoption of artificial intelligence and machine learning in the nuclear industry faces additional barriers, namely regulatory and stakeholder readiness. To overcome these challenges, new solutions must gain regulatory approval and cater to stakeholder needs. The Nuclear Regulatory Committee has a 5-year strategic plan which prepares them for reviewing artificial intelligence technologies in licensee submissions. Early and frequent engagement with the regulator is encouraged. Additionally, artificial intelligence solutions should incorporate human-in-the-loop considerations and offer explainability. Stakeholders must prepare by hiring or training staff to adapt to advancing technology in everyday plant tasks.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Agentic traffic intelligence: Augmented human-in-the-loop scenario generation for microscopic traffic simulation

Traditional microscopic traffic simulation generation often relies on static datasets and manual design, limiting its ability to simulate complex conditions easily. This paper presents a novel framework, Agentic Traffic Intelligence, which combines human approval large language models (LLMs), the Real-Twin tool, and multi-agent systems to perform realistic microscopic traffic simulation scenario generation. The proposed framework incorporates human-in-the-loop (HIL) control, retrieval-augmented generation (RAG), and multi-agent control mechanisms. HIL mechanisms are used to guide multiple LLMs focused on attributes for microscopic simulation generation and to improve the interpretability and transparency of LLM execution for users. RAG enhances context extraction by dynamically integrating external knowledge sources for traffic scenario generation foundations. A multi-agent architecture with supervisory control coordinates the interaction of simulation components, including traffic simulators, control logic, and calibration tools. This enables the synthesis of simulation-ready scenarios that reflect dynamic demand profiles and behavior controls. Furthermore, the framework fuses multisource traffic data with unstructured context and supports iterative refinement through interactive user feedback. Validated through microscopic simulation using Simulation of Urban Mobility, the generated scenarios demonstrate high-fidelity network generation with inflow and turn movement and behavioral calibration, offering a robust and efficient tool for stress-testing and optimizing urban mobility systems.

Hierarchical multi-agent control↗

Scalable in situ non-destructive evaluation of additively manufactured components using process monitoring, sensor fusion, and machine learning

Laser Powder Bed Fusion (L-PBF) Additive Manufacturing (AM) is among the metal 3D printing technologies most broadly adopted by the manufacturing industry. However, the current industry qualification paradigm for critical-application L-PBF parts relies heavily on expensive non-destructive inspection techniques, which significantly limits the use-cases of L-PBF. In situ monitoring of the process promises a less expensive alternative to ex situ testing, but existing sensor technologies and data analysis techniques struggle to detect sub-surface flaws (e.g., porosity and cracking) on production-scale L-PBF printers. In this work, an in situ NDE (INDE) system was engineered to detect subsurface flaws detected in X-Ray Computed Tomography (XCT) directly from process monitoring data. A multilayer, multimodal data input allowed the INDE system to detect numerous subsurface flaws in the size range of 200–1000µm using a novel human-in-the-loop annotation procedure. Furthermore, a framework was established for generating probability-of-detection (POD) and probability-of-false-alarm (PFA) curves compliant with NDE standards by systematically comparing instances of detected subsurface flaws to post-build XCT data. Here, we also introduce for the first time in the AM in situ sensing literature the a 90/95 – the flaw size corresponding to a 90% detection rate on the lower 95% confidence interval of the POD curve. The INDE system successfully demonstrated POD capabilities commensurate with traditional NDE methods. Traditional ML performance metrics were also shown to be inadequate for assessing the ability of the INDE system’s flaw detection performance. It is the hope of the authors that future studies will adopt the POD and PFA approach outlined here to provide better insight into the utility of process monitoring for AM.

36 MATERIALS SCIENCE↗

Load Shedding for Voltage Regulation With Probabilistic Agent Compliance

With the increased observability and controllability of distribution systems, the share of behind-the-meter systems is trending upwards rapidly. As a consequence, the impact of human behaviors on system performance can no longer be ignored and should be reflected in the energy management system models. In this paper, we discuss the problem of distribution system voltage control by active power curtailment where the agent compliance of the load curtailment signal is probabilistic. We discuss the modeling of the optimal voltage control problem with probabilistic agent compliance as a chance-constrained optimization problem, its tractable safe approximation using convex restriction, and a scenario-based mixed-integer reformulation as well as the associated solution method based on augmented Lagrangian method. The numerical simulation on IEEE test system validates the effectiveness of the proposed approach in obtaining high-quality feasible load curtailment signal with low computational cost, which makes it a viable tool for real time decision making.

augmented Lagrangian method↗

Advancements in Development and Testing of Thermal Power Dispatch Simulators

Flexible plant operations and generation (FPOG) offer nuclear power plants (NPPs) the chance to leverage alternative, non-electric revenue streams while ensuring their continued role as reliable, clean, and constant sources of baseload electrical power. The excess thermal energy generated from NPPs during periods of low electricity demand can be channeled as raw materials to numerous industrial processes via a thermal power dispatch (TPD) system. Hydrogen production via high-temperature steam electrolysis (HTSE) is an optimal use case based on technical and economic feasibility. Researchers at Idaho National Laboratory (INL) have conducted previous works that developed and implemented TPD system models within the GSE Solutions Generic Pressurized Water Reactor (GPWR) simulator to support human-in-the-loop (HITL) scenario-based evaluations. The first part of this report documents modifications made to the GPWR TPD model and HMI from the previous iteration in line with a new Sargent and Lundy (S&L) TPD design with an automatic control system. The was done in collaboration with Westinghouse using their three-loop pressurizer water reactor (W3LPWR) simulator which contains an industrial grade automatic control system for the TPD. This was installed in the Human Systems Simulation Laboratory (HSSL) at INL. The second part of the report documents findings from an all-hands-on-deck integration and verification workshop that was conducted in the HSSL over several days. The research team comprised INL human factors and TPD experts, a nuclear engineer from GSE Solutions who implemented the revised TPD model for GPWR, the human-machine interface (HMI) prototyping and human factors team from the University of Idaho, and personnel with operations experience with pressurized water reactors. The workshop provided time and expertise to conduct the final activities to bring the operations, HMI, and simulator into a functional state. The goals of the integration and verification workshop were: 1. to install the revised GPWR TPD model into the HSSL 2. verify the TPD HMI prototype was functional 3. integrate the HTSE Simulink model to GPWR. 3. Issues were identified for resolution, but overall the workshop accomplished its goal to integrated and verify the majority of the intended functional. Future work will resolve the identified issues and use the integrated simulation to support an evaluation and demonstration in the next fiscal year.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Review of low-cost self-driving laboratories in chemistry and materials science: the “frugal twin” concept

This review proposes the concept of a “frugal twin,” similar to a digital twin, but for physical experiments. Frugal twins range from simple toy examples to low-cost surrogates of high-cost research systems. For example, a color-mixing self-driving laboratory (SDL) can serve as a low-cost version of a costly multi-step chemical discovery SDL. Frugal twins already provide hands-on experience for SDLs with low costs and low risks. They can also offer as test beds for software prototyping (e.g., optimization, data infrastructure), and a low barrier to entry for democratizing SDLs. However, there is room for improvement. The true value of frugal twins can be realized in three core areas. Firstly, hardware and software modularity; secondly, purpose-built design (human-inspired vs. hardware-centric vs. human-in-the-loop); and thirdly state-of-the-art (SOTA) software (e.g., multi-fidelity optimization). We also describe the ethical benefits and risks that come with the democratization of science through frugal twins. For future work, we suggest ideas for new frugal twins, SDL educational course outcomes, and a classification scheme for autonomy levels.

36 MATERIALS SCIENCE↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Autonomous Flow Electrochemistry for Accelerated Catalyst Discovery

Our objective is to develop an Autonomous Chemical Experimentation (ACE) platform that accelerates discovery of new catalytic transformations and other energy-relevant chemical reactions and processes. We intentionally designed ACE to be highly modular, both with respect to its rapid deployment to different chemistries and experimental workflows as well as incorporation of a wide range of different AI algorithms. In addition to the development of the core software architecture, initial efforts were made to incorporate Large Language Models to provide human-interpretable reasoning of the optimizer’s actions, and to develop a user-friendly graphical interface for experimental researchers. ACE was demonstrated using a flow electrocatalysis platform containing an inline FTIR spectrometer for real-time analysis and quantification of the reaction outcome. Human-in-the-loop experiments were performed in which a human researcher conducted an experiment using electrode potentials suggested by ACE, then fed the spectral data back to ACE for decision making. After confirming the successful function of the optimizer, efforts were next directed to automation of the hardware and performed full autonomy tests using three reactions: catalytic oxidation of formate, catalytic oxidation of cyclohexanol, and oxidation of hydroquinone. These studies confirm that ACE can close the loop between reaction execution, analysis, and optimization. They also reveal that more improved product detection methods will be essential for ACE to make well-informed decisions for reactions with low conversions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗