Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Centric Artificial Intelligence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data-Centric AI and the Open Energy Data Initiative (OEDI)

This presentation emphasizes the critical importance of data-centric AI. The limitations of model-centric AI when dealing with poor or insufficient data are highlighted, and it is illustrated how training models on inaccurate or noisy data leads to suboptimal results. This talk advocates for a hybrid approach that combines a focus on data quality and model parameters to achieve optimal results. The Open Energy Data Initiative (OEDI) is introduced as a valuable resource for obtaining high-quality energy-related datasets, hosting nearly 2,000 publicly accessible datasets, including 99 solar-related datasets, totaling over 2.7 petabytes of data. OEDI's data lakes enable users to query and work with data without extensive transfers. In conclusion, the significance of data-centric AI and adherence to data curation best practices is emphasized, positioning OEDI as a prime source of high-quality data for AI and machine learning in the renewable energy sector.

AI↗

Explainable Artificial Intelligence Technology for Predictive Maintenance

The domestic nuclear power plant fleet has relied on labor-intensive and time-consuming preventive maintenance programs, thus driving up operation and maintenance costs to achieve high-capacity factors. Artificial intelligence and machine learning can help simplify complex problems, such as diagnosing equipment degradation, to enable more effective decision-making. Benefits will be felt not only within existing analog and digital instrumentation and control, but also work processes, the integration of people with technology, and most importantly, the business case. Together, these hold promise to make nuclear power more efficient and reduce costs associated with operation and maintenance. While the artificial intelligence and machine learning technologies hold significant promise in the nuclear industry, there are challenges or barriers to their adoption. This report outlines the those different machine learning adoption barriers (categorized as historical, technical, economic, regulatory, and user) that the industry must overcome to realize the full benefits of artificial intelligence and machine learning capabilities for long-term economic sustainability. This report also provides solutions for some of these barriers by focusing on improving the explainability of machine learning to encourage trust from the end-user. Trust and explainability are essential to machine learning adoption. This report focuses on research-developed solutions to some of these barriers while analyzing a non-safety-related system, namely the circulating water system. This system frequently experiences waterbox fouling which our models preemptively diagnoses then explains to the operator how those conclusions were reached. This report presents and discusses the inherent trade-off between machine learning performance (in terms of accuracy) and explainability, where highly accurate machine learning methods (such as deep-learning) are the least explainable, and the most explainable methods (such as decision trees) are the least accurate. In addition, explainability of artificial intelligence techniques in terms of transparency and post-hoc metrics are discussed. This report outlines the importance of data novelty and value of new information in evaluating both the explainability and trustworthiness. Novelty detection helps to establish consistency or inconsistency of the new data with respect to the training data. On the other hand, value of information could be a part of the user-centric visualization recommendation system that request additional information to be collected, thereby strengthening the machine learning outcomes. During this project, a copyrighted user-centric visualization that aligns with a human-in-the-loop approach was developed. The user-centric visualization presents different levels of information and can be tailored as per user credentials to gain user confidence. One of the salient features of the user-centric visualization is it presents machine learning methods with explainability metrics. A simplified version of the user-centric visualization was presented to 32 users with varying levels of machine learning expertise. Feedback was solicited to test the hypothesis that the app contained sufficient explainability and that the users would trust the algorithm. Overall, the app was positively received, and the hypothesis was supported. This report discusses the trust-but-verify framework – a potential approach to build user trust artificial intelligence. The framework discusses trust from the human level to artificial intelligence level. The fundamental premise of the trust but verify framework is derived from an observation of nuclear safety culture (i.e., nuclear power plant personnel do not rely on a singular source of data to make a decision). This also ties back to the user-centric visualization that presents different levels of information to achieve both explainability and trustworthiness of artificial intelligence. Even so, the adoption of artificial intelligence and machine learning in the nuclear industry faces additional barriers, namely regulatory and stakeholder readiness. To overcome these challenges, new solutions must gain regulatory approval and cater to stakeholder needs. The Nuclear Regulatory Committee has a 5-year strategic plan which prepares them for reviewing artificial intelligence technologies in licensee submissions. Early and frequent engagement with the regulator is encouraged. Additionally, artificial intelligence solutions should incorporate human-in-the-loop considerations and offer explainability. Stakeholders must prepare by hiring or training staff to adapt to advancing technology in everyday plant tasks.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Roadmap on data-centric materials science

Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

36 MATERIALS SCIENCE↗

Prognostics and Health Management in Nuclear Power Plants: An Updated Method-Centric Review With Special Focus on Data-Driven Methods

In a carbon-constrained world, future uses of nuclear power technologies can contribute to climate change mitigation as the installed electricity generating capacity and range of applications could be much greater and more diverse than with the current plants. To preserve the nuclear industry competitiveness in the global energy market, prognostics and health management (PHM) of plant assets is expected to be important for supporting and sustaining improvements in the economics associated with operating nuclear power plants (NPPs) while maintaining their high availability. Of interest are long-term operation of the legacy fleet to 80 years through subsequent license renewals and economic operation of new builds of either light water reactors or advanced reactor designs. Recent advances in data-driven analysis methods—largely represented by those in artificial intelligence and machine learning—have enhanced applications ranging from robust anomaly detection to automated control and autonomous operation of complex systems. The NPP equipment PHM is one area where the application of these algorithmic advances can significantly improve the ability to perform asset management. This paper provides an updated method-centric review of the full PHM suite in NPPs focusing on data-driven methods and advances since the last major survey article was published in 2015. The main approaches and the state of practice are described, including those for the tasks of data acquisition, condition monitoring, diagnostics, prognostics, and planning and decision-making. Research advances in non-nuclear power applications are also included to assess findings that may be applicable to the nuclear industry, along with the opportunities and challenges when adapting these developments to NPPs. Finally, this paper identifies key research needs in regard to data availability and quality, verification and validation, and uncertainty quantification.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A primer on artificial intelligence in plant digital phenomics: embarking on the data to insights journey

Artificial intelligence (AI) has emerged as a fundamental component of global agricultural research that is poised to impact on many aspects of plant science. In digital phenomics, AI is capable of learning intricate structure and patterns in large datasets. We provide a perspective and primer on AI applications to phenome research. We propose a novel human-centric explainable AI (X-AI) system architecture consisting of data architecture, technology infrastructure, and AI architecture design. We clarify the difference between post hoc models and 'interpretable by design' models. We include guidance for effectively using an interpretable by design model in phenomic analysis. We also provide directions to sources of tools and resources for making data analytics increasingly accessible. In conclusion, this primer is accompanied by an interactive online tutorial.

60 APPLIED LIFE SCIENCES↗

Evaluation of AI-enabled Digital Documented Safety Analysis

The National Reactor Innovation Center (NRIC) is leading a transformative initiative to accelerate advanced reactor deployment by fundamentally reimagining how nuclear safety basis documentation is developed, reviewed, and maintained. Traditional Documented Safety Analysis (DSA) processes for DOE-authorized facilities rely on static, document-centric workflows that consume significant time and resources, exemplified by recent major licensing efforts requiring hundreds of thousands of staff hours and millions of pages of documentation review. These conventional approaches create barriers to the rapid, cost-effective deployment of advanced reactors that America's future energy needs demand. NRIC's DOE Authorization Digital Transformation Project addresses these challenges through an innovative framework that integrates artificial intelligence (AI), digital engineering, and systems-based data management into a cohesive digital ecosystem. This white paper presents NRIC's methodology for evaluating AI-enabled document generation capabilities within this broader digital infrastructure, using the Demonstration of Microreactor Experiments (DOME) facility as a pilot case study. The evaluation will assess an AI tool's ability to generate a Preliminary Documented Safety Analysis (PDSA) through progressive integration stages—from standalone document processing to full digital thread connectivity—while maintaining rigorous verification, validation, and regulatory acceptance standards. By establishing dynamic, traceable connections between design data and safety documentation, NRIC's approach has the potential to reduce both document development time and regulatory review cycles by as much as 50%, while simultaneously improving accuracy, consistency, and traceability. This initiative represents a critical step toward establishing reusable digital infrastructure that reactor developers can leverage to accelerate their path from concept to commercial operation, directly supporting NRIC's mission to demonstrate and deploy advanced nuclear energy technologies.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Ultrafast radiographic imaging and tracking: An overview of instruments, methods, data, and applications

Ultrafast radiographic imaging and tracking (U-RadIT) use state-of-the-art ionizing particle and light sources to experimentally study sub-nanosecond transients or dynamic processes in physics, chemistry, biology, geology, materials science and other fields. These processes are fundamental to modern technologies and applications, such as nuclear fusion energy, advanced manufacturing, communication, and green transportation, which often involve one mole or more atoms and elementary particles, and thus are challenging to compute by using the first principles of quantum physics or other forward models. One of the central problems in U-RadIT is to optimize information yield through, e.g. high-luminosity X-ray and particle sources, efficient imaging and tracking detectors, novel methods to collect data, and large-bandwidth online and offline data processing, regulated by the underlying physics, statistics, and computing power. We review and highlight recent progress in: (a.) Detectors such as high-speed complementary metal-oxide semiconductor (CMOS) cameras, hybrid pixelated array detectors integrated with Timepix4 and other application-specific integrated circuits (ASICs), and digital photon detectors; (b.) U-RadIT modalities such as dynamic phase contrast imaging, dynamic diffractive imaging, and four-dimensional (4D) particle tracking; (c.) U-RadIT data and algorithms such as neural networks and machine learning, and (d.) Applications in ultrafast dynamic material science using XFELs, synchrotrons and laser-driven sources. Hardware-centric approaches to U-RadIT optimization are constrained by detector material properties, low signal-to-noise ratio, high cost and long development cycles of critical hardware components such as ASICs. Interpretation of experimental data, including comparisons with forward models, is frequently hindered by sparse measurements, model and measurement uncertainties, and noise. Alternatively, U-RadIT make increasing use of data science and machine learning algorithms, including experimental implementations of compressed sensing. Machine learning and artificial intelligence approaches, refined by physics and materials information, may also contribute significantly to data interpretation, uncertainty quantification and U-RadIT optimization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

BeyondFingerprinting: AI-guided discovery of robust materials & processes

BeyondFingerprinting was a 2021-2024 Sandia Grand Challenge LDRD exploring the potential to develop new resilient materials and manufacturing processes by taking an artificial-intelligence (AI)-guided approach that integrates human-subject-matter expertise with algorithms enriched with physics-based constraints to unearth process-structure-property correlations. Such algorithms, trained on high-throughput experiments and simulations, are shown to serve as surrogate models that efficiently detect key “fingerprints” in materials data, prognose material performance, and guide effective process improvements. To accelerate broader adoption across mission areas, this AI-guided approach was demonstrated with three complex process-centric exemplars: electroplating, physical vapor deposition, and laser powder bed fusion. Together, these exemplars impact nearly every hardware component relevant to DOE and NNSA national security missions.

36 MATERIALS SCIENCE↗

Scalable Technologies Achieving Risk-Informed Condition-Based Predictive Maintenance Enhancing the Economic Performance of Operating Nuclear Power Plants

The primary objective of the research presented in this report is to develop scalable technologies that are deployable across plant assets and across the nuclear fleet to achieve risk-informed predictive maintenance (PdM) strategies at commercial nuclear power plants (NPPs). Over the years, the nuclear fleet has relied on labor-intensive and time-consuming preventive maintenance (PM) programs, driving up operation and maintenance (O&M) costs to achieve high capacity factors. A well-constructed risk-informed PdM approach for an identified plant asset has been developed in this research, taking advantage of advancements in data analytics, machine learning (ML), artificial intelligence (AI), physics-informed modeling, and visualization. These technologies would allow commercial NPPs to reliably transition from current labor-intensive PM programs to a technology driven PdM program, eliminating unnecessary O&M costs. The work presented in the report is being developed as part of a collaborative research effort between Idaho National Laboratory and Public Service Enterprise Group Nuclear, LLC. This report (1) reflects the results of work by LWRS Program researchers with PSEG, Nuclear LLC-owned Salem and Hope Creek Nuclear Power Plants; (2) presents utilization of circulating water system (CWS) heterogeneous data and fault modes from both the Salem and Hope Creek nuclear power plant sites to develop salient fault signatures associated with each fault mode; (3) describes the integration of component-level predictive models into a robust system-level model enabled by the federated-transfer learning; (4) describes the development of physics-informed model of circulating water pump and motor; (5) develops a scalable risk and economic model; and (6) outlines the development of a user-centric visualization application. The outcomes presented in this report lays the foundation and provides a much-needed technical basis to focus on explainability and trustworthiness of ML and AI-based technologies, as part of future research.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Considerations for Introducing Artificial Intelligence into Nuclear Power Plants

Advanced computational tools and techniques such as artificial intelligence and machine learning (AI/ML) can transform the nuclear power industry. This is necessary given that the economic viability of the existing fleet is in jeopardy and its labor-centric approach to operations and maintenance. Currently, AI/ML research is being undertaken for reactor system design and analysis including fault and accident prognosis, nuclear risk analysis such as plant safety and security evaluation, and plant operations and maintenance including predictive maintenance. Applications include both existing and advanced reactor technologies with the aim of improving operational and business efficiencies. Most every aspect of the organization can benefit, from instrumentation and control, to work planning, to human-machine interactions and business management. AI/ML in nuclear can simplify complex problems and produce more effective decision-making. Nonetheless, careful consideration must be given to the implementation of an AI/ML initiative. The aims of this research are to 1) review barriers to AI/ML adoption within the nuclear power industry, and 2) suggest potential solutions. These barriers are organized along five distinct categories (Figure 1) that are interconnected. The first are historical barriers that track the industry’s development over the decades including worldwide nuclear events that shaped public perceptions. The resulting federal scrutiny and intense safety culture that emerged are discussed. Technical barriers to AI/ML adoption are considerable, and include data privacy concerns, data governance, and the current lack of AI/ML expert knowledge at the plants. The main business case barrier remains cost, but an absence of an industry-wide vision and wide-scale adoption also produces reluctance. Stakeholder readiness is reviewed with special attention given to regulatory readiness. The 5-year strategic plan for AI readiness recently published by the U.S. Nuclear Regulatory Commission is highlighted. Last, adoption barriers at the user level are addressed including the importance of user experience and explainable AI. The AI adoption barriers described here are inter-related and ideally should be addressed in a holistic fashion.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Artificial Intelligence in Nuclear Safeguards; Evaluating Safeguards and Security Risks and Benefits for Advanced and Small Modular Reactor Deployments

Rapidly growing interest in advanced and small modular reactor (A/SMR) technologies presents challenges as well as opportunities for implementing international safeguards and security. A/SMR deployments are expected to be more numerous, more geographically dispersed, and more varied in their designs, placing new demands on the data systems and analytical tools used to support oversight (Alberti et al., 2023; Canadian Nuclear Safety Commission et al., 2024). Because of this variability, the importance and reliance on data systems for A/SMR deployments is expected to be higher than for previous reactor generations. Artificial Intelligence and Machine Learning (AI/ML) offer potential capabilities to address the high variability inherent in A/SMR technology. The beneficiaries of AI-assisted tools include facility operators, government regulators, IAEA inspectors, and A/SMR vendors. This report analyzes how AI/ML-assisted technologies can strengthen the implementation of IAEA safeguards and security measures. It also identifies AI-assisted tools to strengthen operator, facility, and regulator knowledge management practices and examines the potential risks AI/ML-based tools may introduce to IAEA safeguards and security efforts. It concludes with a set of hypothetical, standards-style requirements for AI/ML systems used in safeguards contexts, grounded in an inspector-centric view of system verification. Despite the potential benefits of AI/ML systems, understanding potential intentional and unintentional failure modes is critical for ensuring adequate protection of nuclear materials and facilities. Unique features of A/SMRs including sealed cores, remote and novel paradigms of operation, off-site reactor fabrication, novel fuel forms, and varied refueling requirements, introduce challenges for traditional safeguards technological approaches (Pensado et al., 2024; Federation of American Scientists, 2025). AI/ML systems deployed to address these challenges may introduce new risks requiring systematic evaluation rooted in both AI-specific risk frameworks, such as the NIST AI Risk Management Framework (NIST AI RMF), and established cyber risk management standards such as NIST SP 800-30 (National Institute of Standards and Technology [NIST], 2023; NIST, 2012).

97 MATHEMATICS AND COMPUTING↗

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

Design-Space Data: Informing Common Design Decisions with Pre-Simulated Data

Design Space Exploration (DSE) analysis techniques represent a data-centric approach to integrating performance analysis in early design phases when there is the greatest potential to cheaply improve the energy efficiency of a building. We focus on a novel extension of DSE called Universal Design Space Exploration (UDSE), which leverages massive databases of pre-simulated analysis that represent all possible outcomes of common analysis workflows. These databases, called Design Spaces, become “universal” when a single pre-simulated design space can be re-applied to future unknown projects. Unlike current simulation methods, which require a design to exist before it can be analyzed and often take minutes or hours to simulate, UDSE leverages pre-simulation to deliver rapid and relevant insight as new designs are conceptualized. The data underpinning UDSE enables advanced statistical and Artificial Intelligence methods, allowing UDSE to deliver a greater understanding of the larger problem being explored, rather than simply delivering analysis of several pre-conceived design options. We believe that UDSE can provide instantaneous, relevant analysis for all building design projects at negligible cost. This paper has two main goals, to develop a relevant Universal Design Space that showcases the potential of UDSE and to release this data freely to industry and academia; thereby lowering the barrier to entry to digital literacy in statistics, ML and AI within the architecture, engineering, construction (AEC) industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Machine learning in nuclear materials research

Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes with associated transmutations, high temperature and temperature gradients, mechanical stresses, and corrosive coolants. They also have a wide range of microstructural and chemical makeups, resulting in multifaceted and often out-of-equilibrium interactions. Machine learning (ML) is increasingly being used to tackle these complex time-dependent interactions and aid researchers in developing models and making predictions, sometimes with better accuracy than traditional modeling that focuses on one or two parameters at a time. Conventional practices of acquiring new experimental data in nuclear materials research are often slow and expensive, limiting the opportunity for data-centric ML, but new methods are changing that paradigm. Here we review high-throughput computational and experimental data approaches, especially robotic experimentation and active learning that is based on Gaussian process and Bayesian optimization. We show ML examples in structural materials (e.g., reactor pressure vessel (RPV) alloys and radiation detecting scintillating materials) and highlight new techniques of high-throughput sample preparation and characterizations, and automated radiation/environmental exposures and real-time online diagnostics. Herein, this review suggests that ML models of material constitutive relations in plasticity, damage, and even electronic and optical responses to radiation are likely to become powerful tools as they develop. Finally, we speculate on how the recent trends of using natural language processing (NLP) to aid the collection and analysis of literature data, interpretable artificial intelligence (AI), and the use of streamlined scripting, database, workflow management, and cloud computing platforms that will soon make the utilization of ML techniques as commonplace as the spreadsheet curve-fitting practices of today.

36 MATERIALS SCIENCE↗

Enabling site-specific well leakage risk estimation during geologic carbon sequestration using a modular deep-learning-based wellbore leakage model

Geologic carbon sequestration (GCS) is a promising technology for mitigating net carbon emissions and growing climate concern by storing CO 2 in reservoirs. Oil and gas brownfields are an attractive option for CO 2 storage, but these sites have many historical wellbores from petroleum production and can be a potential leakage pathway for CO 2 or formation brine. Therefore, risk management of GCS operations requires an assessment of potential well leakage. Due to the high uncertainty of the system, stochastic approaches are ideal for quantifying the range of risk behaviors, but they must be computationally efficient in the face of complex physics. Here, we develop a new physics-centric deep learning wellbore model to predict the leakage of CO 2 and brine through leaky wellbores. Multi-physics numerical simulations were used to generate data sets, and physics-informed features were introduced. Neural networks were optimized with an automated searching algorithm. Feature analysis quantifies the impact of each feature on model prediction and confirms the role of physics-inspired parameters. The model shows high predictive performance across a wide range of geologic and injection conditions and well attributes. In conclusion, a case study illustrates how the model is applied to assess well leakage in GCS operations.

58 GEOSCIENCES↗

Demonstration and Evaluation of Explainable and Trustworthy Predictive Technology for Condition-based Maintenance

The domestic nuclear power plant (NPP) fleet has historically relied on labor-intensive and time-consuming predictive maintenance (PdM) programs, thus driving up operation and maintenance (O&M) costs to achieve high-capacity factors. Artificial intelligence (AI) and machine-learning (ML) can help simplify complex problems such as diagnosing equipment degradation to enable more effective decision-making efforts. The benefits of AI will be felt through more efficient plant O&M, improved work processes, and better integration of people and technology. Together, these benefits hold the promise to make nuclear power more sustainable by reducing O&M costs while improving employee engagement. While AI and ML technologies hold significant promise for the nuclear industry, there are challenges or barriers to their adoption. Explainability and trustworthiness of AI are two salient challenges that need to be addressed for wider deployment of these technologies in NPPs. This research focuses specifically on addressing the explainability and trustworthiness of AI technologies to advance the human, technical, and organization (HTO) readiness levels in adopting a risk-informed PdM strategy at commercial NPPs. In addition, this approach can be adapted to enhance the acceptability of AI in other nuclear applications with a few application-specific modifications. The technical approach ensuring wider adoption of AI technologies was developed by Idaho National Laboratory (INL)—in collaboration with Public Service Enterprise Group (PSEG), Nuclear, LLC—by utilizing the circulating water system (CWS) at two PSEG-owned plant sites for demonstration. Focused user studies were performed in collaboration with subject matter experts (SMEs) from PSEG and other nuclear domains to enhance human and organization readiness by building trust in AI-informed technologies. VIsualization for PrEdictive maintenance Recommendation (VIPER)—a Battelle Energy Alliance, LLC, copyrighted software—was developed and expanded to provide a user-centric visualization by incorporating inputs from the collaborating utility, human factors engineering guidelines, and data analysts. The VIPER software enables users, who may be unfamiliar with ML in general, to be interactively engaged by asking technical questions about PdM, work orders, diagnosis results and their confidence levels, the kind of data being used, and the types of ML algorithms employed. This interactive engagement enhances explainability and builds trust. One of the enabling accomplishments was the integration of large language models (LLMs), both text-based and vision-based, in the VIPER software.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗