Digital Documented Safety Analysis: Evaluating the Use of AI for Rapid Generation of Nuclear Safety Basis Documents
Presentation for INL's AI/ML Symposium. Topic: Evaluation of AI-generated documents for nuclear permitting/licensing.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Presentation for INL's AI/ML Symposium. Topic: Evaluation of AI-generated documents for nuclear permitting/licensing.
This paper introduces py-guam, an open-source experimentation framework developed for the NASA Generic Urban Air Mobility simulation (GUAM) environment, facilitating the integration and evaluation of advanced artificial intelligence (AI) algorithms. We present a systems approach which enables the seamless incorporation of data-driven models, including off-nominal and failure state detection, into the GUAM’s Cognitive Architecture (CA). The framework supports customizable experimentation parameters, derives Safety Performance Indicators (SPIs) from UL 4600 safety case analyses, and employs rapid UAM simulations to assess AI impacts on flight performance across diverse scenarios. Through comprehensive testing and validation experiments, we demonstrate GUAM’s capability to enhance safety and efficiency in urban air mobility operations. Additionally, the open-source nature of py-guam fosters community collaboration, ensuring continuous improvement and adaptability to evolving technological advancements. This work establishes a robust tool for developing and testing AI-driven urban air mobility (UAM) systems, advancing the safety and reliability of autonomous urban air vehicles.
Artificial intelligence (AI) has the potential to transform nuclear security operations, offering opportunities to enhance effectiveness while simultaneously introducing new challenges. As AI technologies rapidly evolve, agencies across the United States Government (USG) are researching, implementing, and evaluating various AI models and systems. Given the broad capabilities and applications of these technologies, it is essential for each agency to identify and articulate those areas where it can make meaningful contributions aligned with its mission and expertise. To address this need for strategic focus, in late Fiscal Year 2025 (FY2025), the Office of International Nuclear Security (INS) established an AI Task Force (AITF) to gather input from subject matter experts (SMEs) regarding the most appropriate role INS could serve in researching, evaluating, or implementing AI for nuclear security. The AITF engaged 15 experts from national laboratories with backgrounds in cyber security, physical security, transport security, insider threat mitigation, nuclear engineering, human-systems engineering, and AI/ML development. This white paper summarizes the insights gathered from these SMEs and presents a potential roadmap for INS engagement with AI technologies. The recommendations outlined here are intended to inform INS leadership as they make strategic decisions about resource allocation and program direction in this rapidly evolving technological domain.
BACKGROUND Objective Structured Clinical Evaluations (OSCEs) have long been established as a robust methodology for summative assessment of clinical skills and decision-making during medical education. The recent integration of Artificial Intelligence (AI) into clinical decision-making processes has prompted the need for novel evaluation frameworks to assess the efficacy and reliability of AI clinical decision support system (CDSS) tools. This abstract outlines the process of quantitatively evaluating a novel CDSS (“Doc in a Box” Google 2024) trained on curated medical spaceflight data in the psychomotor domain as it interfaces with a human volunteer acting as the crew medical officer (CMO). PURPOSE The AI CDSS under review was developed as part of the Lunar Command and Control Interoperability (LuCCI) project, which is intended to address a gap in how Lunar Surface Systems (LSS) would interoperate across multiple programs, commercial partners, and international partners. The project objective is to define, prototype, integrate, and evaluate an interoperable lunar command, control, data, and software reference architecture to enable autonomy and informatics capability through common standards across LSS. A multi-modal AI-based CDSS compatible with Federated LSS will assist clinicians in diagnosing and managing complex medical conditions by providing evidence-based recommendations through predictive analytics. Given the critical role of decision-support as NASA continues to evolve its Earth-independent medical operations (EIMO), it is imperative to ensure that such AI tools perform reliably and align with clinical standards during progressive lunar and Martian exploration class missions. METHODS The OSCE framework, traditionally used for evaluating human clinicians, was adapted to assess the AI tool's decision-making capabilities in simulated clinical scenarios. In this adapted OSCE, the AI CDSS was tested across a series of structured clinical scenarios designed to mimic real-life spaceflight patient cases. These scenarios included a range of conditions and complexities, allowing for comprehensive assessment of the tool's performance. Key evaluation metrics included accuracy of diagnosis, timeliness of decision-making, and appropriate recommendations for therapies. The OSCE was scored by human physician evaluators who assessed the AI's recommendations in comparison with expert clinicians' medical decision making to ensure alignment with best practices and the standard of care. RESULTS Preliminary results indicate that the AI CDSS demonstrated high accuracy in diagnostic recommendations and decision support across various scenarios. However, certain limitations were noted, such as occasional discrepancies in handling complex or nuanced cases that required a more contextual understanding. Additionally, the tool scored higher on the diagnostic portion of the rubric, with lower scores in the therapeutic recommendations. These findings highlight the importance of continuous refinement and validation of AI tools through rigorous evaluation frameworks like the OSCE. The adaptation of OSCEs for AI tools presents several advantages, including a structured and reproducible approach to evaluation, the ability to test AI systems in diverse clinical scenarios, and the opportunity to benchmark AI performance against established clinical standards to permit charting of future progress as aerospace medicine evolves as a discipline. Remaining challenges include ensuring that these evaluations capture the full spectrum of clinical decision-making scenarios that will be confronted by CMOs during missions and adequately reflecting real-world variability of the austere spaceflight environment. CONCLUSION Employing OSCEs to evaluate AI clinical decision support tools offers a promising approach to validating their clinical utility and efficacy. This methodology not only provides insights into the tool's performance but also fosters ongoing improvement and alignment with standard of care practices. Future research should focus on refining these evaluation processes and addressing limitations to enhance the integration of AI tools in clinical spaceflight settings. REFERENCES Scott S, Hearns V, Barker MA. Testing Clinical Skills: A Look at the OSCE and USMLE Clinical Skills Exams. S D Med. 2019 Oct;72(10):451-453. Majumder MAA, Kumar A, Krishnamurthy K, Ojeh N, Adams OP, Sa B. An evaluative study of objective structured clinical examination (OSCE): students and examiners perspectives. Adv Med Educ Pract. 2019 Jun 5;10:387-397. Karam VY, Park YS, Tekian A, Youssef N. Evaluating the validity evidence of an OSCE: results from a new medical school. BMC Med Educ. 2018 Dec 20;18(1):313.
The demand for new energy infrastructure is increasing across the United States, but heterogenous permitting processes and embedded requirements across different local jurisdictions can cause project delays, increase “soft costs,” and hinder developer expansion. This study analyzes the variability in local permitting requirements across the U.S. and develops a quantitative approach to describe their clarity and effectiveness in enabling infrastructure project development. By using an Energy Language Model (ELM), a large language model (LLM) for energy technologies, we systematically gathered permitting information from nearly 300 state-, county-, and city-level documents, creating a structured dataset of requirements and procedures on an unprecedented scale and speed. Our analysis revealed that local (city and county) permitting requirement documents are underrepresented compared to state-level guidance documents, which can impede timely and cost-effective installation of new electric infrastructure. Our validation process showed that the final database has an accuracy of approximately 95%. We, further, created a new quantitative method to score permitting requirements for clarity and efficiency, with electric vehicle supply equipment as an initial use case. The average local permitting document scored a 1.8 out of 5, which we interpret as meaning that half of the requirements developers face when installing electric infrastructure are ambiguous, increasing both cost and time. We also created a “Generalized Permit Process”, highlighting common procedural steps and identifying specific opportunities for municipalities to improve their documentation. This research establishes a systematic and scalable framework for evaluating the complexities of local infrastructure permitting processes by combining LLM-powered data collection and quantitative scoring. The framework enables policymakers and developers to identify and mitigate procedural bottlenecks, with the expectation that these improvements can accelerate application review and approval, reduce project costs, and expedite connection to utility distribution grids. As a foundational approach for streamlining local project development processes, this study’s methods are intended to be extended to a wide range of energy applications.
The aging water distribution system in the United States, constructed mainly during the 1970s with some pipes dating back 125 years, is experiencing significant deterioration leading to substantial water losses. Along with the potential for water loss savings, improvements in the distribution system by using leak detection technologies can create net energy and cost savings. In this work, a new framework has been presented to calculate the economic level of leakage within water supply and distribution systems for two primary leak detection technologies (acoustic vs. satellite). In this work, a new framework is presented to calculate the economic level of leakage (ELL) within water supply and distribution systems to support smart infrastructure in smart cities. A case study focused using water audit data from Atlanta, Georgia, compared the costs of two leak mitigation technologies: conventional acoustic leak detection and artificial intelligence–assisted satellite leak detection technology, which employs machine learning algorithms to identify potential leak signatures from satellite imagery. The ELL results revealed that conducting one survey would be optimum for an acoustic survey, whereas the method suggested that it would be expensive to utilize satellite-based leak detection technology. However, results for cumulative financial analysis over a 3-year period for both technologies revealed both to be economically favorable with conventional acoustic leak detection technology generating higher net economic benefits of USD 2.4 million, surpassing satellite detection by 50%. A broader national analysis was conducted to explore the potential benefits of US water infrastructure mirroring the exemplary conditions of Germany and The Netherlands. Achieving similar infrastructure leakage index (ILI) values could result in annual cost savings of $\$4$–$\$4.8$ billion and primary energy savings of 1.6–1.9 TWh. These results demonstrate the value of combining economic modeling with advanced leak detection technologies to support sustainable, cost-efficient water infrastructure strategies in urban environments, contributing to more sustainable smart living outcomes.
The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER?s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use. PowerPoint for conference that was reviewed in PRS and LRS PRS/CON-25-05379 and INL/CON-25-82946
The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER’s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use.
Large language models (LLMs) are increasingly capable of answering technical questions, synthesizing domain knowledge, and supporting engineering workflows. For nuclear science and engineering, these capabilities require careful, domain-specific evaluation before they can be credibly incorporated into safety-related activities, regulatory review, or technical decision support. This paper presents preliminary results from benchmarking framework for evaluating LLM capabilities in nuclear contexts. The framework is organized into three evaluation categories: nuclear fundamentals, general dual-use knowledge, and plant specific knowledge. These categories are intended to distinguish general nuclear engineering competence from broader technical reasoning and more context-dependent nuclear knowledge. Initial evaluations focus on nuclear fundamentals using questions representative of the knowledge expected of a nuclear professional engineer. Results indicate that contemporary frontier models perform at a high level and substantially exceed the performance of older model generations, with some models approaching saturation of the current benchmark. These findings suggest both the rapid improvement of LLM capabilities in specialized technical domains and the need for more discriminating evaluation methods. The paper presents the benchmark structure, preliminary model-comparison results, and ongoing work. This work supports development of verifiable, responsible, and safety-conscious methods for assessing AI systems in nuclear engineering applications.
Benchmarks provide a standardized method for evaluating different AI models, enabling reproducibility and comparison between models, and facilitating scientific progress. As AI models continue to develop rapidly, incorporating new datasets, capabilities, and architectures becomes more complicated. Therefore, the current static benchmarks become increasingly irrelevant. The MLCommons team argues that to make AI benchmarks more relevant, it involves making the benchmarks themselves more dynamic, as well as technical innovations that make it easier for scientists and researchers at all levels to use and contribute to the benchmarks. The current progress in technical innovation is a software that allows for a detailed view of a collection of AI benchmarks to be output in various formats that are easily readable and accessible.
An evaluation is made of the application of novel, AI-capabilities-related technologies to aerospace systems. Attention is given to expert-system shells for Space Shuttle Orbiter mission control, manpower and processing cost reductions at the NASA Kennedy Space Center's 'firing rooms' for liftoff monitoring, the automation of planetary exploration systems such as semiautonomous mobile robots, and AI for battlefield staff-related functions.
Abstract Cryo-electron microscopy (cryo-EM) has revolutionized structural biology by enabling the determination of high-resolution 3-Dimensional (3D) structures of large biological macromolecules. Protein particle picking, the process of identifying individual protein particles in cryo-EM micrographs for building protein structures, has progressed from manual and template-based methods to sophisticated artificial intelligence (AI)-driven approaches in recent years. This review critically examines the evolution and current state of cryo-EM particle picking methods, with an emphasis on the impact of AI. We conducted a comparative evaluation of popular AI-based particle picking methods, using both general machine learning metrics and specific cryo-EM structure determination metrics. This analysis involved constructing the 3D density map from the picked protein particles and assessing the obtained resolution and particle orientation diversity, underscoring the significant impact of AI on cryo-EM particle picking. Despite the advancements, we also identified key obstacles, such as handling complex micrographs with small proteins. The analysis provides insights into the future development of more sophisticated and fully automated AI methods in cryo-EM particle recognition.
This report is concerned with the application of software quality and evaluation measures to AI software and, more broadly, with the question of quality assurance for AI software. Considered are not only the metrics that attempt to measure some aspect of software quality, but also the methodologies and techniques (such as systematic testing) that attempt to improve some dimension of quality, without necessarily quantifying the extent of the improvement. The report is divided into three parts Part 1 reviews existing software quality measures, i.e., those that have been developed for, and applied to, conventional software. Part 2 considers the characteristics of AI software, the applicability and potential utility of measures and techniques identified in the first part, and reviews those few methods developed specifically for AI software. Part 3 presents an assessment and recommendations for the further exploration of this important area.
Explore the source record for details and available documents.
Airborne Imaging Spectrometer-2 (AIS-2) data was acquired over two paired conifer stands for the purpose of detecting differences in spectral reflectance between stressed and natural canopies. Water stress was induced in a stand of Norway spruce and white pine by severing the sapwood near the ground. Water stress during the AIS flights was evaluated through shoot water potential and relative water content measurements. Preliminary analysis with raw AIS-2 data using SPAM indicates that there were small, inconsistent differences in absolute spectral reflectance in the near infrared 0.97 to 1.3 micron between the stressed and natural canopies.
Researchers at NASA/GSFC evaluated various automated inspection systems (AIS) technologies using test boards with known defects in surface mount solder joints. These boards were complex and included almost every type of surface mount device typical of critical assemblies used for space flight applications: X-ray radiography; X-ray laminography; Ultrasonic Imaging; Optical Imaging; Laser Imaging; and Infrared Inspection. Vendors, representative of the different technologies, inspected the test boards with their particular machine. The results of the evaluation showed limitations of AIS. Furthermore, none of the AIS technologies evaluated proved to meet all of the inspection criteria for use in high-reliability applications. It was found that certain inspection systems could supplement but not replace manual inspection for low-volume, high-reliability, surface mount solder joints.
- Near real-time identification of airborne dust in satellite imagery is important for mitigating the adverse effects of dust storms on human activities. - False color Red-Green-Blue (RGB) imagery has been used for dust detection, but it has limitations and can be difficult to interpret. - The NASA Short-term Research and Transition (SPoRT) center has developed a night-time dust detection random forest (NT-DustTracker-AI, Berndt et al. 2021) model using NASA/NOAA Geostationary Operational Environmental Satellite-16 (GOES-16) Advanced Baseline Imager (ABI) infrared imagery as inputs. - The SPoRT center has partnered with the NOAA National Weather Service to evaluate the model for use in weather forecasting operations, and preliminary results have been positive. - The SPoRT center has expanded the model to cover both day and night, continuing to use infrared imagery as inputs. This new model, known as DustTracker-AI, has shown good agreement with available dust observations
The US team of the European led "MIcrostructure Formation in CASTing of Technical Alloys under Diffusive and Magnetically Controlled Convective Conditions" (MICAST) program recently received a third Aluminum - 7wt% silicon alloy that was processed in the microgravity environment aboard the International Space Station. The sample, designated MICAST#2-12, was directionally solidified in the Solidification with Quench Furnace (SQF) at a constant rate of 40micometers/s through an imposed temperature gradient of 31K/cm. Procedures taken to evaluate the state of the sample prior to sectioning for metallographic analysis are reviewed and rational for measuring the microstructural constituents, in particular the primary dendrite arm spacing (Lambda (sub1)), is given. The data are presented, put in context with the earlier samples, and evaluated in view of a relevant theoretical model.