Search NASASearch

SEARCH · Search NASA

Results for “Heuristic Evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Heuristic Evaluation Methods Applied to a Predictive Maintenance Chatbot

The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER’s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use.

99 - GENERAL AND MISCELLANEOUS

Heuristic Evaluation Methods Applied to a Predictive Maintenance Chatbot

The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER?s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use. PowerPoint for conference that was reviewed in PRS and LRS PRS/CON-25-05379 and INL/CON-25-82946

99 - GENERAL AND MISCELLANEOUS

HPDF Data Catalog and Lakehouse Demo Heuristic Evaluation and Roadshow Feedback Reports

The December 2025 roadshow gathered rapid feedback from community members on the proof of concept AmSC/HPDF data lakehouse and catalog solution 1. This provided the opportunity to demonstrate our current state of progress and gather input on workflows and AI agents. We recognize that the proofs of concept interfaces (OpenMetadata and Goose) are not intended for our end users so we have captured feedback to keep in mind as additional technical and conceptual design work is undertaken.

97 MATHEMATICS AND COMPUTING

Improving the User Interface of the DeepLynx Data Warehouse

DeepLynx is an open-source ontology-based data warehouse created by INL to support the creation and life cycle of digital engineering projects, with a particular emphasis on digital twins [1]. Digital twins are systems that represent physical assets and process in a real-time digital environment [1]. Most well-known commercial data warehouses use Graphical User Interfaces (GUIs) for users to interact with their systems [3]. Limited publications have addressed the design of these interfaces and understanding of their target users. The current users and development team acknowledge the need to improve the current UI, not just for aesthetics but to improve functionality and workflow of DeepLynx. Traditional data warehouse users are developers, data scientists and business analysts [2]. DeepLynx users have a vast range of experience using data warehouses, and diverse roles, including engineers, scientists and management positions. Because there is a broader audience of target users for DeepLynx than a typical data warehouse, it is essential that DeepLynx has a useable and intuitive user interface. To achieve this the team performed human-computer interaction methods, including a Heuristic Evaluation of current UI using Neilsen’s Usability Heuristic, create personas based on current users by designing a user survey, data analysis and develop of personas. Followed by a redesign of the UI following using Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design in industry standard software Figma. Lastly a Heuristic Evaluation of new UI design, using Neilsen’s Usability Heuristic and User testing of redesign UI and have a group of users complete a Thinking Aloud Test of the new UI. Preliminary results of the Heuristic Evaluation of current UI arise issue with Consistency and Standards, Visibility of System Status, Match System and Real World and Recognition Rather than Recall. These issues were addressed in the proposed redesign by applying Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design. Next steps include formalized list of lessons learned and design implications for future publications.

97 MATHEMATICS AND COMPUTING

Characterizing and communicating uncertainty: lessons from NASA’s Carbon Monitoring System

Navigating uncertainty is a critical challenge in all fields of science, especially when translating knowledge into real-world policies or management decisions. However, the wide variance in concepts and definitions of uncertainty across scientific fields hinders effective communication. As a microcosm of diverse fields within Earth Science, NASA’s Carbon Monitoring System (CMS) provides a useful crucible in which to identify cross-cutting concepts of uncertainty. The CMS convened the Uncertainty Working Group (UWG), a group of specialists across disciplines, to evaluate and synthesize efforts to characterize uncertainty in CMS projects. This paper represents efforts by the UWG to build a heuristic framework designed to evaluate data products and communicate uncertainty to both scientific and non-scientific end users. We consider four pillars of uncertainty: origins, severity, stochasticity versus incomplete knowledge, and spatial and temporal autocorrelation. Using a common vocabulary and a generalized workflow, the framework introduces a graphical heuristic accompanied by a narrative, exemplified through contrasting case studies. Envisioned as a versatile tool, this framework provides clarity in reporting uncertainty, guiding users and tempering expectations. Beyond CMS, it stands as a simple yet powerful means to communicate uncertainty across diverse scientific communities.

54 ENVIRONMENTAL SCIENCES

Faster solutions to the interdiction defense problem using suboptimal solutions

The interdiction defense (ID) problem solves a defender-attacker-defender model where the defender and attacker share the same set of components to harden and target. Here, we build upon the best response intersection (BRI) algorithm by developing the BRI with suboptimal solutions (BRI-SS) algorithm to solve the ID problem. The BRI-SS algorithm utilizes off-the-shelf optimization solvers that return suboptimal solutions at no additional computation cost. We derive novel cuts from suboptimal solutions, reducing the number of iterations required for the algorithm to converge while maintaining optimality guarantees. We also present a heuristic that utilizes all obtained suboptimal solutions to select the next defense to evaluate at each iteration. We perform computational experiments applied to power grid interdiction on standard test cases. Our results demonstrate that the BRI-SS algorithm consistently outperforms the BRI algorithm across all test cases.

Computer science

NLR Core Modeling & Decision Support Capabilities: FASTSim, RouteE, T3CO & OpenPATH

This project is part of the program area to develop and improve core capabilities for the Energy-Efficient Mobility Systems (EEMS) program that enable research, development and deployment of advanced mobility solutions and enhance the EEMS Program's ability to address system-level transportation challenges. Advancements to the Future Automotive Systems Technology Simulator (FASTSim), Route Energy Prediction Model (RouteE), Transportation Technology Total Cost of Ownership (T3CO) and Open Platform for Agile Trip Heuristics (OpenPATH) core capabilities under this project supports the overall EEMS Program goals to effectively evaluate energy and mobility impacts of future transportation technologies and services, and to identify the most promising pathways to reduce transportation costs and environmental harms, and to improve mobility access. This presentation was prepared for the 2026 Annual Merit Review of this project.

33 ADVANCED PROPULSION SYSTEMS

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

kilic, Ozgur Ozan [Brookhaven National Laboratory

Resource Assessment for Distributed Wind Energy: An Evaluation of Best-Practice Methods in the Continental US

Current wind resources within the United States (US) indicate a potential to profitably install nearly 1,400 gigawatts of distributed wind (DW) capacity. This amount is equivalent to over half of the United States’ current energy demand from electricity, making it enough to power millions of homes and businesses and replace countless fossil fuel-based generating plants. Despite the potential growth of DW in the US, deployments are presently hindered by a lack of confidence in resource estimation methods. One potential challenge is that smaller-scale turbines, with hub heights of 40 meters or less, are disproportionately impacted by obstacles such as buildings and vegetation. These obstacles may produce complex wake effects, best modeled with high-fidelity complex fluid dynamics (CFD) models that are too computationally expensive to use for routine siting and resource assessment. Thus, installers today make use of heuristics and simple equations to approximate the impact of obstacles while also leveraging long-term resource data from commercial or publicly available atmospheric models. This study evaluates these historical and commonly used methods alongside new lower-order obstacle models produced from CFD simulations and measurement-based bias correction. The preliminary results from this study show the importance of taking care in the choice and application of mesoscale atmospheric models and the significant value of bias correction using measurements from nearby meteorological towers. Detailed obstacle modeling provides only modest additional gains in performance and, in some cases, can add error, especially at sites where turbines have already been located to avoid obvious impact from upwind obstacles. These findings reinforce the importance of collecting in situ measurements and suggest that obstacle models may be better applied in practice to automated or computer-aided siting, rather than in economic wind resource assessments.

17 WIND ENERGY

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC

Data-Driven Voltage Regulation of Distribution Grid Using Nonlinear Autoregressive Model with Exogenous Inputs (NARX)

This article proposes data-driven control via a nonlinear autoregressive model with exogenous inputs (NARX) for real-time voltage regulation of a modified feeder using reactive power sources. Traditional voltage control strategies rely on rule-based heuristics or optimization techniques, which often require detailed system models and extensive computational resources. The NARX-based controller learns system dynamics from historical data and predicts optimal reactive power dispatch in real-time for voltage correction. The proposed approach is evaluated on a power system feeder model under varying load and network conditions. Simulation results demonstrate that the NARX-based controller achieves improved voltage regulation, offering higher adaptability to system fluctuations. This study highlights the potential of data-driven control for enhancing the reliability of power distribution networks.

Donge, Vrushabh [ORNL] (ORCID:0000000306062803)

Conformalized-KANs: Uncertainty Quantification with Coverage Guarantees for Kolmogorov-Arnold Networks (KANs) in Scientific Machine Learning

This paper explores uncertainty quantification (UQ) methods in the context of Kolmogorov–Arnold Networks (KANs). We apply an ensemble approach to KANs to obtain a heuristic measure of UQ, enhancing interpretability and robustness in modeling complex functions. Building on this, we introduce Conformalized-KANs, which integrate conformal prediction, a distribution-free UQ technique, with KAN ensembles to generate calibrated prediction intervals with guaranteed coverage.} Extensive numerical experiments are conducted to evaluate the effectiveness of these methods, focusing particularly on the robustness and accuracy of the prediction intervals under various hyperparameter settings. We show that the conformal KAN predictions can be applied to recent extensions of KANs, including Finite Basis KANs (FBKANs) and multifideilty KANs (MFKANs). The results demonstrate the potential of our approaches to significantly improve the reliability and applicability of KANs in scientific machine learning.

• Artificial intelligence (AI) / machine learning

“Understanding Robustness Lottery”: A Geometric Visual Comparative Analysis of Neural Network Pruning Approaches

Deep learning approaches have provided state-of-the-art performance in many applications by relying on large and overparameterized neural networks. However, such networks are very brittle and are difficult to deploy on resource-limited platforms. Model pruning, i.e., reducing the size of the network, is a widely adopted strategy that can lead to a more robust and compact model. Many heuristics exist for model pruning, but our understanding of the pruning process remains limited due to the black-box nature of a neural network model. Empirical studies show that some heuristics improve performance whereas others can make models more brittle. Here, this work aims to shed light on how different pruning methods alter the network’s internal feature representation and the corresponding impact on model performance. To facilitate a comprehensive comparison and characterization of the high-dimensional model feature space, we introduce a visual geometric analysis of feature representations. We evaluated a set of critical geometric concepts decomposed from the commonly adopted classification loss and used them to design a visualization system to compare and highlight the impact of pruning on model performance and feature representation. The proposed tool provides an environment for an in-depth comparison of pruning methods and a comprehensive understanding of how the model responds to common data corruption. By leveraging the proposed visualization, machine learning researchers can reveal the similarities between pruning methods and redundancy in robustness evaluation benchmarks, obtain geometric insights about the differences between pruned models that achieve superior robustness performance, and identify samples that are robust or fragile to model pruning and common data corruption.

Li, Zhimin [Univ. of Utah, Salt Lake City, UT (Uni

Scan2Sim: Software to Convert Network Scans to Emulations

Within operational technology (OT) systems design, the construction of testing environments for simulation is often a tedious, manual process that slows down safety and security evaluations. This document details the design and functionality of Scan2Sim, a program designed to construct high-fidelity topological schematics for OT systems without significant manual human input. Scan2Sim may take as input a detailed network scan of a system, and produces an instruction set to re-create the original scanned network within a virtualized simulation network. This construction is achieved via heuristic methods of machine template selection, which allows for a fast, performant approach to automated environment construction. The current tool is designed to produce topology schematics compatible with the Minimega, a tool designed by Sandia National Laboratories for repeatable experimentation management.

97 MATHEMATICS AND COMPUTING

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI

CHEMREASONER: Heuristic Search over a Large Language Model’s Knowledge Space using Quantum-Chemical Feedback

The discovery of new catalysts is essential for the design of new and more efficient chemical processes in order to transition to a sustainable future. We introduce an AI-guided computational screening framework unifying linguistic reasoning with quantum-chemistry based feedback from 3D atomistic representations. Our approach formulates catalyst discovery as an uncertain environment where an agent actively searches for highly effective catalysts via the iterative combination of large language model (LLM)-derived hypotheses and atomistic graph neural network (GNN)-derived feedback. Identified catalysts in intermediate search steps undergo structural evaluation based on spatial orientation, reaction pathways, and stability. Scoring functions based on adsorption energies and barriers steer the exploration in the LLM's knowledge space toward energetically favorable, high-efficiency catalysts. We introduce planning methods that automatically guide the exploration without human input, providing competitive performance against expert-enumerated chemical descriptor-based implementations. By integrating language-guided reasoning with computational chemistry feedback, our work pioneers AI-accelerated, trustworthy catalyst discovery.

artificial intelligence

Evaluating the impacts of Variable Message Signs on Airport Curbside Performance Using Microsimulation

Curbs play a vital role in facilitating vehicle access and egress for individuals at airports. Inefficiently allocating this resource hinders airport accessibility and productivity, resulting in congestion, longer travel times, and increased pollution. As airport demand fluctuates throughout the day and grows over time, airports face intensified curbside pressure. Yet, curb management research is significantly less robust at airports than in urban areas. Given the unbalanced nature of airport demand—riders tend to arrive simultaneously at specific entrances at certain hours—Variable Message Sign (VMS) arises as a cost-effective technology to divert vehicles from congested to underutilized curbs. Still, VMS implementation faces a significant challenge. Historically, airports have managed VMS heuristically and by intuition rather than an evidence-based approach. This research investigates the impacts of implementing VMS on curb performance at airports. By considering different driver compliance rates (DCR), we aim to determine when the sign should be turned on and off to diverge traffic to avoid undesired externalities while enhancing curb performance. Using a validated agent-based microsimulation model, VISSIM, we analyzed the Seattle-Tacoma (SeaTac) Airport as a case study. We modeled sixteen VMS management scenarios and a baseline where the message sign is not displayed, diverging vehicles between the departures and arrivals access levels at four different moments (early morning, morning, afternoon, and late night). We quantified the effects of VMS using seven metrics, including curb productivity index (CPI), curb accessibility (CA), queue length, queue duration, delay, vehicle counts, and emissions. The results of each scenario were compared against the baseline using absolute and relative changes and Repeated Measures ANOVA. Overall, VMS improved curb performance and traffic conditions at the airport, reducing emissions by 14.8% to 8.9%. Moreover, significant reductions in queue length (1,150 ft to 100 ft) and duration (15 to 144 minutes) were observed in the sending link under all VMS policies. However, impacts on the receiving link varied based on congestion, with significant increases in queue duration (9.8 to 24 min) when congested but no substantial changes in free flow. Notably, diverging vehicles to congested links resulted in non-significant results, and activating late and deactivating late VMS affected curb productivity (-5.8% to -61.4%), curb accessibility (-16.5% to -25.8%), cumulative counts (-33.4% to -59.4%), and vehicle delay (95.98% to 594.3%). Activating VMS before congestion begins in the sending link and deactivating before a queue forms in the receiving link yield the most significant improvements: 8.1% to 10.1% in CPI, 9.4% to9.6% in CA, -29.3% to -77.9% in total delay, -11.6% to -13.9% in total emissions, and 101% to 103% in cumulative counts. As the analysis was made with a wide range of time periods, access levels, driver compliance rates, and scenarios, we believe our findings can provide valuable insights into how airports should manage VMS. Our work introduces a novel approach to the scientific airport literature, as some of our metrics were previously unexplored. Additionally, we propose a methodology that other airports can adopt to maximize their curb performance.

Gutierrez, Jorge D.