Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI Risks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Forward for the new edition of: “What Every Engineer Should Know About Risk Engineering and Management"

In aerospace and engineering in general, the major metrics are capability, cost, and safety/risk. With the increasingly rapid emergence and utilization of new technologies and systems of increasing complexity, ensuring safety becomes more difficult. The technology and practice/applications of safety/risk technologies require updating in concert with capability and systems technology changes. Hence the new, updated edition of this risk engineering book. Major changes in technology and applications, now and going forward, increasingly involve artificial intelligence (AI)/autonomy and complex systems, which are rapidly developing and moving targets when it comes to risk/safety analysis. These introduce both new risks and safety issues, and they alter more usual ones. The current reality is that often the best AI is when it is “Black Boxed”, developed “independently”, without detailed human understanding of how and by what processes decisions and results are produced. There are ongoing efforts to make the AI processes more understandable by humans, with results to be determined. Also, trusted and true autonomy is free of human intervention, which would require machines to ideate to solve in real time issues that arise due to unknown unknowns and even known unknowns. Machines using generative adversarial network (GANs) and other approaches are beginning to ideate. Overall, the capabilities and practice of AI/autonomy utilization is a work in progress with the rate of progress substantial and the impacts upon system risk/safety major.

Dennis M. Bushnell↗

Demonstration and Evaluation of Explainable and Trustworthy Predictive Technology for Condition-based Maintenance

The domestic nuclear power plant (NPP) fleet has historically relied on labor-intensive and time-consuming predictive maintenance (PdM) programs, thus driving up operation and maintenance (O&M) costs to achieve high-capacity factors. Artificial intelligence (AI) and machine-learning (ML) can help simplify complex problems such as diagnosing equipment degradation to enable more effective decision-making efforts. The benefits of AI will be felt through more efficient plant O&M, improved work processes, and better integration of people and technology. Together, these benefits hold the promise to make nuclear power more sustainable by reducing O&M costs while improving employee engagement. While AI and ML technologies hold significant promise for the nuclear industry, there are challenges or barriers to their adoption. Explainability and trustworthiness of AI are two salient challenges that need to be addressed for wider deployment of these technologies in NPPs. This research focuses specifically on addressing the explainability and trustworthiness of AI technologies to advance the human, technical, and organization (HTO) readiness levels in adopting a risk-informed PdM strategy at commercial NPPs. In addition, this approach can be adapted to enhance the acceptability of AI in other nuclear applications with a few application-specific modifications. The technical approach ensuring wider adoption of AI technologies was developed by Idaho National Laboratory (INL)—in collaboration with Public Service Enterprise Group (PSEG), Nuclear, LLC—by utilizing the circulating water system (CWS) at two PSEG-owned plant sites for demonstration. Focused user studies were performed in collaboration with subject matter experts (SMEs) from PSEG and other nuclear domains to enhance human and organization readiness by building trust in AI-informed technologies. VIsualization for PrEdictive maintenance Recommendation (VIPER)—a Battelle Energy Alliance, LLC, copyrighted software—was developed and expanded to provide a user-centric visualization by incorporating inputs from the collaborating utility, human factors engineering guidelines, and data analysts. The VIPER software enables users, who may be unfamiliar with ML in general, to be interactively engaged by asking technical questions about PdM, work orders, diagnosis results and their confidence levels, the kind of data being used, and the types of ML algorithms employed. This interactive engagement enhances explainability and builds trust. One of the enabling accomplishments was the integration of large language models (LLMs), both text-based and vision-based, in the VIPER software.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Responsible Adoption of Artificial Intelligence (AI) in Electric Grid Operations

The future of the grid will be powered by AI—or undermined by it. Artificial intelligence is rapidly reshaping grid operations, improving fault detection, forecasting accuracy, and real-time optimization. As AI systems move closer to operational decision loops, however, they introduce new consequence pathways: expanded attack surfaces, model integrity risks, regulatory exposure, and human-automation challenges. This talk presents a consequence-driven framework for deploying AI responsibly in the electric grid. Attendees will gain practical strategies to strengthen resilience, boost reliability, and deploy AI securely — ensuring the grid of the future is not only smarter but safer.

25 - ENERGY STORAGE↗

Data Centers and Digital Assurance Workshop 3 – Mitigations for Digital Assurance Risks

The third session of the TADA (Technical Assistance for Digital Assurance) Data Centers Cohort, held on November 18, 2025, focused on developing mitigation strategies for digital assurance risks identified in previous workshops. Hosted by Idaho National Laboratory (INL) and ScottMadden, the session emphasized the application of Cyber-Informed Engineering (CIE) to data center infrastructure, particularly at the utility–data center interface. Participants revisited and ranked key digital assurance risks, including architecture and interface weaknesses, governance gaps, and AI-enabled threats. The workshop introduced the 12 principles of CIE, advocating for consequence-focused design, engineered controls, and secure information architecture to proactively reduce cyber-physical vulnerabilities. These principles were applied to critical data center systems such as power distribution, UPS, cooling, SCADA/BMS, and grid-forming batteries. The session also addressed governance challenges at the interconnection boundary, highlighting the need for clear roles in telemetry sharing, firmware management, and trip settings. Special attention was given to emerging risks from behind-the-meter (BTM) generation, including reverse-power flow and the integration of small modular reactors (SMRs), which shift data centers from large loads to complex generation nodes. Participants explored how interconnection agreements can serve as enforceable instruments for digital assurance, and reviewed gaps in current standards such as NERC CIP, IEC 62443, and IEEE 1547. The workshop concluded with pathways to standardization, including model agreement language, state-level programs, and expanded NERC guidance. INL also presented tools and frameworks for secure procurement and supplier risk management, reinforcing the need for integrated engineering and policy solutions to secure the evolving data center–grid ecosystem. Session 3 of 3.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Earth Independent Medical Operations (EIMO) DATASCOPE Technical Interchange Meeting 21st August 2023: Background and Summary of Discussion

An aspiration for EIMO datascope is to realize artificial intelligence-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A vision proposed to the meeting participants was that of a “system of systems,” whereby EIMO will utilize AI-supported natural language processing and machine learning techniques to synthesize embedded reference databases and real-time data streams [input vectors] from multiple data sources to continuously and seamlessly assess crew health & performance. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will ideally have a degree of mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats.

Artificial Intelligence↗

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

J Lemery↗

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

Medical Operations↗

Universal Electronic‐Structure Relationship Governing Intrinsic Magnetic Properties in Permanent Magnets

An electronic-structure-centered perspective is presented on permanent-magnet (PM) design, highlighting two key levers, that is, saturation magnetization (M s ), governed by 3d-band filling and exchange physics, and magnetocrystalline anisotropy energy (MAE), arising from spin-orbit coupling (SOC) on anisotropic orbital populations. Reviewing current practices, including DFT-based MAE/J ij extraction, atomistic-spin and micromagnetic modeling, and high-throughput machine learning (ML) pipelines, three bottlenecks limiting predictive discovery is identified that is i) electronic-structure accuracy for small MAE (sensitive to functional choice, Hubbard U, and many-body effects), ii) finite-temperature and kinetic realism (phonon/magnon renormalization, ordering kinetics), and iii) descriptor and multiscale decoupling (lack of SOC-weighted and orbital-resolved fingerprints). Deep dives into the electronic-structure of Nd─Fe─B and Fe─N show how these fingerprints govern magnetic performance, motivating DFT- and quantum-mechanics-based descriptors for discovery. Unbiased, structure-driven exploration, coupled with high-throughput simulations, ML, generative AI, and reasoning models, accelerates candidate identification and propagates insights across scales. Addressing supply-chain risks, on future needs of designing “critical-element-free” magnets with tailored microstructure and high energy products is emphasized. By integrating electronic fingerprints, AI reasoning, and multiscale modeling, a practical roadmap is provided for rare-earth-lean or rare-earth free, high-performance, sustainable PMs.

Singh, Prashant [Ames Laboratory, and Iowa State U↗

JWST Pathfinder Telescope Integration

The James Webb Space Telescope (JWST) is a 6.5m, segmented, IR telescope that will explore the first light of the universe after the big bang. In 2014, a major risk reduction effort related to the Alignment, Integration, and Test (AI&T) of the segmented telescope was completed. The Pathfinder telescope includes two Primary Mirror Segment Assemblies (PMSA's) and the Secondary Mirror Assembly (SMA) onto a flight-like composite telescope backplane. This pathfinder allowed the JWST team to assess the alignment process and to better understand the various error sources that need to be accommodated in the flight build. The successful completion of the Pathfinder Telescope provides a final integration roadmap for the flight operations that will start in August 2015.

JWST↗

Genesis Mission-Enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE)

Argonne National Laboratory is supporting the U.S. Department of Transportation’s (USDOT’s) Bureau of Transportation Statistics (BTS) with collaborative research on development and application of privacy preserving AI frameworks that leverage unmatched AI expertise and secure computing resources made available through the U.S. Genesis Mission1 . This research advances U.S. energy security goals by supporting a safe offshore energy industry with secure, domain-specific AI tools to analyze confidential industry datasets collected by BTS to rapidly improve identification of hazards, precursors, and systemic safety risks in high-risk operational environments. The staged, security-first approach begins with development and testing of Argonne’s Genesis Mission-enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE) framework within Argonne’s accredited secure computing enclave (ABLE) leveraging Argonne’s AI scientific assistant substrate (AISAC). Methods to build synthetic datasets were developed together with BTS for use in preparing synthetic datasets that can be used to validate data containment, governance, and security controls in the ABLE environment. Future research directions would focus on applying the Genesis-SAFE framework to CIPSEA-protected datasets entirely within ABLE to support confidentiality-preserving analysis of safety risks, trends, and contributing factors.

Kim, Hyekyung [Argonne National Laboratory (ANL), ↗

Risks Associated with Sharing the MOSSAIC APIs

The MOSSAIC APIs contain two files which, in theory, could be used to discover information about the pathology report data from the SEER registries on which the AI models were trained. In this document, we explain the contents of these files and assess the associated risk. APPENDIX A contains a set of slides to aid in the dissemination of this information.

97 MATHEMATICS AND COMPUTING↗

Human Factors in Space Exploration

The exploration of space is one of the most fascinating domains to study from a human factors perspective. Like other complex work domains such as aviation (Pritchett and Kim, 2008), air traffic management (Durso and Manning, 2008), health care (Morrow, North, and Wickens, 2006), homeland security (Cooke and Winner, 2008), and vehicle control (Lee, 2006), space exploration is a large-scale sociotechnical work domain characterized by complexity, dynamism, uncertainty, and risk in real-time operational contexts (Perrow, 1999; Woods et ai, 1994). Nearly the entire gamut of human factors issues - for example, human-automation interaction (Sheridan and Parasuraman, 2006), telerobotics, display and control design (Smith, Bennett, and Stone, 2006), usability, anthropometry (Chaffin, 2008), biomechanics (Marras and Radwin, 2006), safety engineering, emergency operations, maintenance human factors, situation awareness (Tenney and Pew, 2006), crew resource management (Salas et aI., 2006), methods for cognitive work analysis (Bisantz and Roth, 2008) and the like -- are applicable to astronauts, mission control, operational medicine, Space Shuttle manufacturing and assembly operations, and space suit designers as they are in other work domains (e.g., Bloomberg, 2003; Bos et al, 2006; Brooks and Ince, 1992; Casler and Cook, 1999; Jones, 1994; McCurdy et ai, 2006; Neerincx et aI., 2006; Olofinboba and Dorneich, 2005; Patterson, Watts-Perotti and Woods, 1999; Patterson and Woods, 2001; Seagull et ai, 2007; Sierhuis, Clancey and Sims, 2002). The human exploration of space also has unique challenges of particular interest to human factors research and practice. This chapter provides an overview of those issues and reports on sorne of the latest research results as well as the latest challenges still facing the field.

FROM↗

Artificial Intelligence (AI) Methods for Augmenting the IMPACT Tool Evidence Library

Development of the Evidence Library for use with the IMPACT probability risk assessment tool took several years and involved a staggering amount of effort from a multi-disciplinary team. A very significant amount of the labor effort to collect, assess and finalize the Clinical Finding Form (CliFF) for each of the 119 medical conditions was provided by physician subject matter experts from the Exploration Medical Capability (ExMC) Element Clinical and Science Team. Many AI tools such as ChatGPT are excellent at summarizing large amounts of information and the current project was initiated to determine how such tools might streamline laborious processes, e.g., review and summarization of many scientific research publications, to execute key steps more efficiently in the process of developing CliFFs. The process for collecting the evidence which is found in the CliFFs is well documented in the Evidence Library Methods document (ELM; HRP-48036*). Using ELM and the CliFF development instructions as a guideline, a team of developers is leveraging Microsoft Azure AI tools and services along with open-source frameworks, to construct an AI-assisted automated pipeline. This pipeline is designed to search, retrieve, and process the necessary data sources, and ultimately help generate the final version of a CliFF. Currently, the large language model evaluates the relevance of each source material to spaceflights, either as direct evidence or as an analog. Additionally, the model assists in extracting keywords and generating brief summaries to enhance augmented retrieval and search processes in later stages of CliFF development. Once the data is ready, the model can perform semantic search and retrieval, generating and extracting valuable information for the CliFF. For instance, it can handle epidemiological statistical data, such as incidence rates and the likelihood of best or worst-case scenarios. The steps that required reading and summarizing articles were viewed as providing the greatest return on investment since large language models are very efficient and accurate in summarizing large amounts of text. Since labor effort to complete the original CliFF was not recorded with sufficient granularity, comparisons with an AI tool-generated CliFF will provide merely an approximation of time saved. Upon completion of the process, the CliFF for the medical condition “appendicitis” generated with the support of AI-based methods will serve as a proof-of-concept and will be compared to the original appendicitis CliFF to determine if use of the tools resulted in content and conclusory similarity. Based upon the results from face validation of the two CliFFs, modifications to the process will be made if necessary and additional condition CliFFs will be evaluated. Ultimately, CliFFs for the entire set of medical conditions will be created with the assistance of AI tools. Depending on the cost savings realized, CliFFs for additional medical conditions can be created to expand the Evidence Library. Future direction includes specifying the characteristics of the reviewer (prompting the AI tools to generate output assuming the reviewer is a sub-specialist physician, or nurse or EMT/medic) to determine if the effects on AI-generated output are different based on knowledge, skills and abilities. *Exploration Medical Capability Evidence Library Methods, HRP-48036 Rev A, July 2022.

Ali Al↗

Machine-Learning-Based Mapping and Modeling of Solar Energy with Ultra-High Spatiotemporal Granularity

Despite the rapid growth of solar energy, we still lack a dynamic, high-fidelity database that tracks the spatiotemporal variations of solar PVs and their associated infrastructures across different places at a spatially resolved scale. The absence of such data presents a barrier to various applications such as solar PV growth projection, solar energy integration, solar incentive design, and climate risk assessment. In this project, we aim to bridge this gap by developing AI-based algorithms to extract granular information about solar PV installations and their associated infrastructures (i.e., distribution grids) from widely available unstructured data like remote sensing images and street views. As a result, we have built the Solar Energy Atlas, a fine-grained, large-scale geospatial overlay of distributed solar PVs and distribution grids. On top of it, we have advanced the understanding of solar adoption and distribution grid vulnerability to climate-induced extremes. Our major contributions can be summarized as follow: (1) By developing new AI algorithms, we have built the most comprehensive solar PV spatiotemporal database covering the entire US. This is the first time we obtained the exact GPS locations, size, subtype, and installation year information for rooftop solar PVs across the US. This database can be used for solar PV growth projection, solar energy integration, solar energy policy analysis and design, and spatially-resolved climate risk assessment. (2) Leveraging this database, we have uncovered the socioeconomic driving factors that are correlated with earlier onset of solar adoption and higher saturated adoption levels. We have identified the heterogeneity in the effects of different types of financial incentives on solar adoption and provided implications for tailoring incentive design based on local income levels to promote equitable solar adoption. (3) We have developed a distribution grid GIS mapping algorithm which can obtain granular geospatial and topology information about distribution grids using multi-modal open data, reducing the dependency on hard-to-obtain smart meter data of conventional approaches. It shows effectiveness in both the U.S. and Sub-Saharan Africa. Using this algorithm, we have uncovered the non-uniform vulnerability of distribution grids to wildfires in California in the aspects of undergrounding protection and Distributed Energy Resources (DER) preparedness. This has provided important implications for improving the affordability and equity of grid adaptation approaches. (3) We have made our produced database publicly available and provided user-friendly interface to enable various stakeholders and the general public to interact with the data. We have also integrated the produced data into the Data Commons platform to enable the public to access the data and correlate it with other location-specific characteristics simply using natural language as queries. The impact of our project is three-fold: (1) New algorithms for mapping solar PVs and distribution grids across space and time, which are open source to facilitate researchers and industry; (2) New databases of solar PVs and distribution grids that have been made publicly available for engineering, social, and policy applications; (3) New understandings and actionable insights on the potential approaches to promoting solar adoption and reducing energy infrastructure vulnerabilities. In this report, we start by discussing the project background and motivation (section 5), followed by the overview of project objectives (section 6). Results and discussion for each task are presented in section 7. Significant accomplishments are summarized in section 8. This report will be concluded by discussing the paths forwards (section 9), products (section 10), and team roles (section 11).

14 SOLAR ENERGY↗

REFSafE: A RAG-Enabled Framework for Predictive Risk Analysis and Automated Safety Report Generation in Mission-Critical Environments

Operational safety in mission-critical environments requires AI systems that are accurate, interpretable, and resistant to hallucination. We present an agentic Retrieval-Augmented Generation (RAG) framework, REFSafe, for grounded hazard analysis and automated safety report generation. The system integrates Large Language Models (LLMs) with structured operational data, historical incident repositories, policy documents, and external authoritative sources. Through iterative agentic reasoning, the framework retrieves, verifies, and synthesizes evidence prior to generation, enforcing citation-backed outputs with explicit source attribution (documents, links, and prior events) to ensure traceability and trust. To mitigate hallucinations and unsupported claims, all risk assessments and forecasts are constrained to retrieved evidence, with confidence signals derived from retrieval relevance and source consistency. A transparent pipeline enables subject matter experts (SMEs) to validate predictions, and provide structured feedback, forming a continuous performance calibration loop. Preliminary deployment demonstrates improved reliability in hazard detection and safety/vulnerability report generation. This work advances trustworthy, evidence-grounded AI for predictive safety intelligence in mission-critical operations.

Das, Sanjay [ORNL] (ORCID:0009000542591915)↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Evaluation of off-road terrain with static stereo and monoscopic displays

The National Aeronautics and Space Administration is currently funding research into the design of a Mars rover vehicle. This unmanned rover will be used to explore a number of scientific and geologic sites on the Martian surface. Since the rover can not be driven from Earth in real-time, due to lengthy communication time delays, a locomotion strategy that optimizes vehicle range and minimizes potential risk must be developed. In order to assess the degree of on-board artificial intelligence (AI) required for a rover to carry out its' mission, researchers conducted an experiment to define a no AI baseline. In the experiment 24 subjects, divided into stereo and monoscopic groups, were shown video snapshots of four terrain scenes. The subjects' task was to choose a suitable path for the vehicle through each of the four scenes. Paths were scored based on distance travelled and hazard avoidance. Study results are presented with respect to: (1) risk versus range; (2) stereo versus monocular video; (3) vehicle camera height; and (4) camera field-of-view.

Yorchak, John P.↗

Trustworthiness and Trust: Identifying Factors that Drive Successful Human-AI Interaction in Nuclear Power Plant Applications

Emerging technologies such as artificial intelligence (AI) and machine learning (ML) are rapidly evolving and considered a promising tool for efficient and continued safe operations of the U.S. nuclear power plants (NPPs). Emerging AI techniques like large language models (LLMs) are one such technology that may support personnel at existing NPPs perform work more efficiently. For example, operators may query the current operational status of a power plant via a chat interface leveraging LLMs to access plant-related information in an interactive manner rather than manually collecting various sensor data for tasks such as surveillances or completing work orders. This is a fundamental shift in the way operators currently perform their tasks today. The literature of human-automation interaction indicates that trust is a crucial factor that drives successful interaction between a human operator and an automated system, like an AI-infused NPP application. This work presents the results of a literature review on key factors that relate to trust in AI/LLM technologies for NPP applications. The relevant literature of human factors and cognitive engineering has identified various factors related to trust including trustworthiness, performance characteristics, operator skill and perceived risk. This preliminary literature review will guide development and evaluation of models involving the identified factors influencing trust in AI and develop a framework for human-centered design for interface between humans and AI. By addressing trust, this work supports developing a technical basis for designing key characteristics of AI/LLM to support calibrated trust, which will ultimately support wide-scale adoption of AI/LLM technologies, as well as ensure safe, effective, and reliable use.

99 - GENERAL AND MISCELLANEOUS↗