Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI Evaluations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Energy and AI: Evaluating Future Grid and Water Stress Due to Data Centers

Projections of the need for new data centers to support Artificial Intelligence (AI) are large but highly uncertain. Recent projections indicate up to a 15% annual growth rate in data center electricity demand within the next 5-10 years. Given that most electric utilities are required to have a reserve margin of roughly the same magnitude as the projected growth in demand, these new data center loads could soon threaten resource adequacy and reliability unless data centers build their own generation, interruptible loads are negotiated, commensurate new capacity and/or transmission is built, or some combination of these options. Similarly, depending on the cooling technology and geographic location of new data centers, they could threaten water adequacy in water scarce regions. This presentation highlights the grid and water implications of new data centers to support AI.

Mongird, Kendall (ORCID:0000000328077088)↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

An Evaluation of AI Models’ Performance for Three Geothermal Sites

Current artificial intelligence (AI) applications in geothermal exploration are tailored to specific geothermal sites, limiting their transferability and broader applicability. This study aims to develop a globally applicable and transferable geothermal AI model to empower the exploration of geothermal resources. This study presents a methodology for adopting geothermal AI that utilizes known indicators of geothermal areas, including mineral markers, land surface temperature (LST), and faults. The proposed methodology involves a comparative analysis of three distinct geothermal sites—Brady, Desert Peak, and Coso. The research plan includes self-testing to understand the unique characteristics of each site, followed by dependent and independent tests to assess cross-compatibility and model transferability. The results indicate that Desert Peak and Coso geothermal sites are cross-compatible due to their similar geothermal characteristics, allowing the AI model to be transferable between these sites. However, Brady is found to be incompatible with both Desert Peak and Coso. The geothermal AI model developed in this study demonstrates the potential for transferability and applicability to other geothermal sites with similar characteristics, enhancing the efficiency and effectiveness of geothermal resource exploration. This advancement in geothermal AI modeling can significantly contribute to the global expansion of geothermal energy, supporting sustainable energy goals.

Energy & Fuels↗

Ranking and Classifying AI Benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece C. [Cornell U.]↗

Classifying and rating AI benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark’s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece [Fermilab]↗

Construction of approximate invariants for nonintegrable Hamiltonian systems

We present a method to construct high-order polynomial approximate invariants (AI) for nonintegrable Hamiltonian dynamical systems and apply it to a modern ring-based particle accelerator. Taking advantage of a special property of one-turn transformation maps expressed as square matrices, AIs can be constructed order by order iteratively. Evaluating AI with simulation data, we observe that AI’s fluctuation is actually a measure of chaos. Through minimizing the fluctuations, the stable region of long-term motions, i.e., the dynamic aperture of the accelerator, could be enlarged.

36 MATERIALS SCIENCE↗

Construction of approximate invariants for non-integrable Hamiltonian systems

We present a method to construct high-order polynomial approximate invariants (AI) for non integrable Hamiltonian dynamical systems, and apply it to a modern ring-based particle accelerator. Taking advantage of a special property of one-turn transformation maps in the form of a square matrix, AIs can be constructed order-by-order iteratively. Evaluating AI with simulation data, we observe that AI’s fluctuation is actually a measure of chaos. Through minimizing the fluctuations, the stable region of long-term motions, i.e., the dynamic aperture of the accelerator, could be enlarged.

43 PARTICLE ACCELERATORS↗

Poster: Responsible Adoption of Artificial Intelligence (AI) in Electric Grid Operations

The rapid integration of artificial intelligence (AI) in the utility transmission and distribution (T&D) sector is revolutionizing traditional grid management practices. As utilities encounter complexities from evolving consumer behaviors and energy integration, AI becomes a critical solution for enhancing grid monitoring, fault detection, and operational optimization. However, increased reliance on interconnected technologies introduces significant cybersecurity risks, regulatory compliance challenges, and human factors concerns. This study proposes a strategic, responsible and consequence-driven approach to AI implementation, examining the dual nature of AI adoption by highlighting its transformative benefits for utilities and associated risks. It provides utilities with a framework for evaluating AI integration, enabling them to navigate challenges and capitalize on opportunities to achieve greater reliability, efficiency, and resilience in an increasingly complex energy landscape.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Evaluation of AI-enabled Digital Documented Safety Analysis

The National Reactor Innovation Center (NRIC) is leading a transformative initiative to accelerate advanced reactor deployment by fundamentally reimagining how nuclear safety basis documentation is developed, reviewed, and maintained. Traditional Documented Safety Analysis (DSA) processes for DOE-authorized facilities rely on static, document-centric workflows that consume significant time and resources, exemplified by recent major licensing efforts requiring hundreds of thousands of staff hours and millions of pages of documentation review. These conventional approaches create barriers to the rapid, cost-effective deployment of advanced reactors that America's future energy needs demand. NRIC's DOE Authorization Digital Transformation Project addresses these challenges through an innovative framework that integrates artificial intelligence (AI), digital engineering, and systems-based data management into a cohesive digital ecosystem. This white paper presents NRIC's methodology for evaluating AI-enabled document generation capabilities within this broader digital infrastructure, using the Demonstration of Microreactor Experiments (DOME) facility as a pilot case study. The evaluation will assess an AI tool's ability to generate a Preliminary Documented Safety Analysis (PDSA) through progressive integration stages—from standalone document processing to full digital thread connectivity—while maintaining rigorous verification, validation, and regulatory acceptance standards. By establishing dynamic, traceable connections between design data and safety documentation, NRIC's approach has the potential to reduce both document development time and regulatory review cycles by as much as 50%, while simultaneously improving accuracy, consistency, and traceability. This initiative represents a critical step toward establishing reusable digital infrastructure that reactor developers can leverage to accelerate their path from concept to commercial operation, directly supporting NRIC's mission to demonstrate and deploy advanced nuclear energy technologies.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluation of AI-Enabled Digital Documented Safety Analysis: A Case Study

Safety basis documentation development and review under U.S. Department of Energy (DOE) authorization have emerged as critical constraint throttling deployment of advanced nuclear reactors, with traditional processes demanding extraordinary resource investment that delays the delivery of these technologies. Traditional Documented Safety Analysis (DSA) processes rely on static documents with limited traceability [U.S. DOE]. The regulatory review and engagement processes are similarly constrained, often requiring significant effort and extensive manual verification. The scale of this challenge is exemplified by the U.S. Nuclear Regulatory Commission (NRC) review of the NuScale application, which required over 250,000 staff hours and the evaluation of approximately two million pages of documentation [Bergman 2021]. The volume and complexity of information within nuclear licensing applications or authorization reviews demands innovative approaches to document generation and data management.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Artificial Intelligence for (AI) Nuclear Security: Expert Perspectives on AI Priorities for the Office of International Nuclear Security

Artificial intelligence (AI) has the potential to transform nuclear security operations, offering opportunities to enhance effectiveness while simultaneously introducing new challenges. As AI technologies rapidly evolve, agencies across the United States Government (USG) are researching, implementing, and evaluating various AI models and systems. Given the broad capabilities and applications of these technologies, it is essential for each agency to identify and articulate those areas where it can make meaningful contributions aligned with its mission and expertise. To address this need for strategic focus, in late Fiscal Year 2025 (FY2025), the Office of International Nuclear Security (INS) established an AI Task Force (AITF) to gather input from subject matter experts (SMEs) regarding the most appropriate role INS could serve in researching, evaluating, or implementing AI for nuclear security. The AITF engaged 15 experts from national laboratories with backgrounds in cyber security, physical security, transport security, insider threat mitigation, nuclear engineering, human-systems engineering, and AI/ML development. This white paper summarizes the insights gathered from these SMEs and presents a potential roadmap for INS engagement with AI technologies. The recommendations outlined here are intended to inform INS leadership as they make strategic decisions about resource allocation and program direction in this rapidly evolving technological domain.

97 MATHEMATICS AND COMPUTING↗

An AI-driven framework for evaluating local and state authorities’ permitting processes

The demand for new energy infrastructure is increasing across the United States, but heterogenous permitting processes and embedded requirements across different local jurisdictions can cause project delays, increase “soft costs,” and hinder developer expansion. This study analyzes the variability in local permitting requirements across the U.S. and develops a quantitative approach to describe their clarity and effectiveness in enabling infrastructure project development. By using an Energy Language Model (ELM), a large language model (LLM) for energy technologies, we systematically gathered permitting information from nearly 300 state-, county-, and city-level documents, creating a structured dataset of requirements and procedures on an unprecedented scale and speed. Our analysis revealed that local (city and county) permitting requirement documents are underrepresented compared to state-level guidance documents, which can impede timely and cost-effective installation of new electric infrastructure. Our validation process showed that the final database has an accuracy of approximately 95%. We, further, created a new quantitative method to score permitting requirements for clarity and efficiency, with electric vehicle supply equipment as an initial use case. The average local permitting document scored a 1.8 out of 5, which we interpret as meaning that half of the requirements developers face when installing electric infrastructure are ambiguous, increasing both cost and time. We also created a “Generalized Permit Process”, highlighting common procedural steps and identifying specific opportunities for municipalities to improve their documentation. This research establishes a systematic and scalable framework for evaluating the complexities of local infrastructure permitting processes by combining LLM-powered data collection and quantitative scoring. The framework enables policymakers and developers to identify and mitigate procedural bottlenecks, with the expectation that these improvements can accelerate application review and approval, reduce project costs, and expedite connection to utility distribution grids. As a foundational approach for streamlining local project development processes, this study’s methods are intended to be extended to a wide range of energy applications.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evaluating Acoustic vs. AI-Based Satellite Leak Detection in Aging US Water Infrastructure: A Cost and Energy Savings Analysis

The aging water distribution system in the United States, constructed mainly during the 1970s with some pipes dating back 125 years, is experiencing significant deterioration leading to substantial water losses. Along with the potential for water loss savings, improvements in the distribution system by using leak detection technologies can create net energy and cost savings. In this work, a new framework has been presented to calculate the economic level of leakage within water supply and distribution systems for two primary leak detection technologies (acoustic vs. satellite). In this work, a new framework is presented to calculate the economic level of leakage (ELL) within water supply and distribution systems to support smart infrastructure in smart cities. A case study focused using water audit data from Atlanta, Georgia, compared the costs of two leak mitigation technologies: conventional acoustic leak detection and artificial intelligence–assisted satellite leak detection technology, which employs machine learning algorithms to identify potential leak signatures from satellite imagery. The ELL results revealed that conducting one survey would be optimum for an acoustic survey, whereas the method suggested that it would be expensive to utilize satellite-based leak detection technology. However, results for cumulative financial analysis over a 3-year period for both technologies revealed both to be economically favorable with conventional acoustic leak detection technology generating higher net economic benefits of USD 2.4 million, surpassing satellite detection by 50%. A broader national analysis was conducted to explore the potential benefits of US water infrastructure mirroring the exemplary conditions of Germany and The Netherlands. Achieving similar infrastructure leakage index (ILI) values could result in annual cost savings of $\$4$–$\$4.8$ billion and primary energy savings of 1.6–1.9 TWh. These results demonstrate the value of combining economic modeling with advanced leak detection technologies to support sustainable, cost-efficient water infrastructure strategies in urban environments, contributing to more sustainable smart living outcomes.

acoustic leak detection↗

Heuristic Evaluation Methods Applied to a Predictive Maintenance Chatbot

The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER?s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use. PowerPoint for conference that was reviewed in PRS and LRS PRS/CON-25-05379 and INL/CON-25-82946

99 - GENERAL AND MISCELLANEOUS↗

Heuristic Evaluation Methods Applied to a Predictive Maintenance Chatbot

The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER’s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use.

99 - GENERAL AND MISCELLANEOUS↗