Search NASASearch

SEARCH · Search NASA

Results for “Document Generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluation of AI-enabled Digital Documented Safety Analysis

The National Reactor Innovation Center (NRIC) is leading a transformative initiative to accelerate advanced reactor deployment by fundamentally reimagining how nuclear safety basis documentation is developed, reviewed, and maintained. Traditional Documented Safety Analysis (DSA) processes for DOE-authorized facilities rely on static, document-centric workflows that consume significant time and resources, exemplified by recent major licensing efforts requiring hundreds of thousands of staff hours and millions of pages of documentation review. These conventional approaches create barriers to the rapid, cost-effective deployment of advanced reactors that America's future energy needs demand. NRIC's DOE Authorization Digital Transformation Project addresses these challenges through an innovative framework that integrates artificial intelligence (AI), digital engineering, and systems-based data management into a cohesive digital ecosystem. This white paper presents NRIC's methodology for evaluating AI-enabled document generation capabilities within this broader digital infrastructure, using the Demonstration of Microreactor Experiments (DOME) facility as a pilot case study. The evaluation will assess an AI tool's ability to generate a Preliminary Documented Safety Analysis (PDSA) through progressive integration stages—from standalone document processing to full digital thread connectivity—while maintaining rigorous verification, validation, and regulatory acceptance standards. By establishing dynamic, traceable connections between design data and safety documentation, NRIC's approach has the potential to reduce both document development time and regulatory review cycles by as much as 50%, while simultaneously improving accuracy, consistency, and traceability. This initiative represents a critical step toward establishing reusable digital infrastructure that reactor developers can leverage to accelerate their path from concept to commercial operation, directly supporting NRIC's mission to demonstrate and deploy advanced nuclear energy technologies.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Privacy-Aware RAG-Enabled LLMs for Collaborative AI in Organizations

Recent advancements in Large Language Models (LLMs) based on Transformer architectures have significantly improved capabilities in natural language processing and generation. However, deploying LLMs for inter-organizational communication poses challenges, in ensuring privacy and facilitating effective collaboration. This paper introduces a novel decentralized inference meta-agent chatbot that leverages privacy-aware Retrieval-Augmented Generation (RAG)-enabled LLMs for collaborative AI communication across organizations. Built on Microsoft’s Autogen, the platform enables LLMs to autonomously refine responses, enhancing accuracy and relevance. It incorporates advanced hallucination mitigation techniques using Uptrain and a privacy-focused RAG framework that employs synthetic document generation to protect sensitive information. Comprehensive evaluations demonstrate the platform’s effectiveness in maintaining contextual relevance and stringent privacy standards, effectively addressing critical challenges in LLM-enhanced collaborative AI communication. This work represents a significant step toward secure and efficient inter-organizational collaboration using advanced generative AI technologies.

97 - MATHEMATICS AND COMPUTING

Framework for a Digital Documented Safety Analysis

This framework is developed to progress the digital implementation of digital tools applied to the DOE authorization process, with future applications to NRC SAR development/review, to accelerate the design and review processes of advanced nuclear reactors. The engineering design and licensing process for nuclear reactors is currently burdened by a document-based approach that leads to duplications and errors due to a lack of traceability among numerous static documents. Changes to design information require labor-intensive manual tracing through these documents, creating a high potential for human error. The adoption of a digital ecosystem, utilizing a digital thread to link various aspects of project design and analysis, promises dynamic documentation generation, automatic updates, and error reduction. Model Based Definition (MBD) and Product Lifecycle Management (PLM) tools are central to this digital transformation.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Evaluation of AI-Enabled Digital Documented Safety Analysis: A Case Study

Safety basis documentation development and review under U.S. Department of Energy (DOE) authorization have emerged as critical constraint throttling deployment of advanced nuclear reactors, with traditional processes demanding extraordinary resource investment that delays the delivery of these technologies. Traditional Documented Safety Analysis (DSA) processes rely on static documents with limited traceability [U.S. DOE]. The regulatory review and engagement processes are similarly constrained, often requiring significant effort and extensive manual verification. The scale of this challenge is exemplified by the U.S. Nuclear Regulatory Commission (NRC) review of the NuScale application, which required over 250,000 staff hours and the evaluation of approximately two million pages of documentation [Bergman 2021]. The volume and complexity of information within nuclear licensing applications or authorization reviews demands innovative approaches to document generation and data management.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Total Effective Dose from Radiologic Emissions from INL Facilities for Calculation of Population Dose for the INL 2024 Annual Site Environmental Report

Total effective radiation dose from airborne releases was calculated using air dispersion modeling performed by the National Oceanic and Atmospheric Administration (NOAA) Idaho Falls Office using their HYSPLIT computer model (Stein et al. 2015; Draxler et al. 2013), and the Dose Multi-Media (DOSEMM) dose assessment model (Rood 2019) . The objective of these calculations was to provide a grid of total effective dose across a model domain that encompasses a 50-mile (80-km) radius from any Idaho National Laboratory (INL) Site source. In addition to INL Site sources, releases from the Radiological and Environmental Sciences Laboratory (RESL) (IF-683) and IF-603 located at the INL Research Center (IRC) within the Idaho Falls city limits were also included. Due to tracking limitations, radionuclides released from IF-611 and IF-603 are modeled as released from IF-603. The dose results will be combined with GIS software to compute a total population dose for the calendar year (CY) 2024 and will be reported in the INL Annual Site Environmental Report (ASER). This report does not cover the population dose calculation and only documents generation of the gridded dose file.

42 - ENGINEERING

LEBT Buncher / Feed Forward / Frequency Generation: System Design Document (SDD)

The LAMP Low Energy Beam Transport (LEBT) transfers a continuous beam at 100 keV from the ion source to RFQ in LEBT. A LEBT buncher imposes an energy tilt to initiate velocy bunching in the chopped beam pulse about 25 ns long to form a short MPEG bunch. One possible option for the LEBT buncher based on a two-gap LC-circuit driven structure was considered, which is similar to the existing LANSCE low-frequncy buncher (LFB). Another possible option is a non-resonant element driven by a pulse-forming-network that provides a single pulse at the repetition frequency of MPEG beam, with a period of 1.8 µs. The LEBT operation is synchronized with the accelerator timing system.

43 PARTICLE ACCELERATORS

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL),

Dual Context: Leveraging Structured Application Context for Code Generation and Runtime Feature Activation via Chat Interfaces

Integrating artificial intelligence (AI) capabilities into software applications typically involves two common paths. For developers, AI assists in generating and documenting source code and other related software engineering efforts. For users, AI assists them through question-and-answer exchanges via chatbots. Both approaches have their value, but neither effectively leverages the modularity of component-based architectures that modern web application frameworks offer. We implement a proof of concept within a centralized suite of applications used for the Atmospheric Radiation Measurement (ARM) Data Center Operational Tools, where we introduce a third integration path through the ARM Context Engine (ACE). ACE is a context driven system that uses structured contextual specifications to enable Large Language Models (LLMs) to render interactive and feature-rich user interface (UI) components directly within chat responses, alongside or in place of conventional text outputs. These specifications serve two important purposes across what we call code context and UI context. Code context provides AI-assisted development tools with structured application knowledge beyond raw code, including component relationships, architectural patterns and schematic information, enabling the generation of consistent, well-structured code. UI context defines the rules for enabling and rendering component features at runtime based on the user's natural language input, allowing end users to activate capabilities such as data export, filtering, and pagination within chat responses, without requiring code changes or redeployment. We demonstrate, through a comparative evaluation against general-purpose AI chatbots, that context-driven component rendering provides interactive capabilities that text-based responses cannot replicate, including deterministic component behavior, application-consistent design language, and on-demand feature activation. A development effort comparison further shows that features that traditionally require multi-step development cycles can be activated with a single naturallanguage request. In this ongoing work, we present ACE as an emerging approach to AI integration that positions modular, well-documented software architecture as the foundation for AI-ready applications. ACE treats context as a shared resource across both development and user-facing AI, bringing cohesion to conventionally disconnected efforts, bridging developer tooling and end-user capabilities within a single framework.

Tadimeti, Vijay [ORNL]

Liquid Abrasive Cutter Power Generation Unit Requirements

This document details the requirements for the purchase of a power generating unit (PGU) for an updated Liquid Abrasive Cutter (LAC) for use in field conditions. The system will provide high pressure water and will be integrated with existing waterjet cutting attachments to cut various materials. The system will be standalone, requiring only water, and electrical energy. Cutting heads, attachments, and a garnet delivery system are not included in this specification. This system is an evolution of the existing LAC (developed at LLNL). The PGU specified herein will have the ability to operate on electrical power via an electric motor, and accept input power from a variety of sources, such as an engine-driven generator, battery pack, solar, shore power, or micro-grid.

42 ENGINEERING

Generative Artificial Intelligence Tools for Red Teams

This document analyzes the role of Generative Artificial Intelligence (GenAI) tools in cybersecurity, particularly for red teaming. While GenAI accelerates initial security assessments, its effectiveness wanes with complexity, necessitating experienced assessors. The review critiques marketing claims, highlights ethical concerns regarding uncensored models for cybercrime, and advocates for a robust defense strategy supported by skilled professionals.

97 MATHEMATICS AND COMPUTING

Deep Cyber-Physical Situational Awareness for Energy Systems: A Secure Foundation for Next-Generation Energy Management

This document provides the final report for the CYPRES project. The purpose is (1) to highlight and summarize its major accomplishments and (2) to provide guidance on how its outcomes have informed and can inform important additional research and technology transfer. The goal of CYPRES was the research, development, and demonstration of a security-oriented next generation cyber-physical EMS for electric power systems that detects malicious and abnormal events through the fusion of cyber and physical data. To achieve this, the CYPRES project team researched, developed, and built a prototype of the solution, referred to as the CYPRES EMS. The CYPRES EMS is a proof-of-concept cyber-physical platform that demonstrates the management of the energy system, communications, security, and cyber-physical grid modeling and analytics. As part of the capabilities of the CYPRES EMS, the team designed and developed a suite of power system applications for monitoring, risk analyses, detection, and control that are inherently cyberaware. At its core, the project aimed to research, develop, and demonstrate a security-oriented next-generation cyber-physical Energy Management System (EMS) capable of detecting malicious and abnormal events through the innovative fusion of cyber and physical data. This approach represents a fundamental shift from traditional EMS, reimagining how critical infrastructure can be protected through unified cyber-aware and physics-aware secure data flow pipelines. The project’s cornerstone deliverable, the CYPRES EMS, serves as a proof-of-concept cyber-physical platform that revolutionizes the management of energy systems, communications, security, and cyber-physical grid modeling and analytics. This prototype implements a comprehensive suite of power system applications for monitoring, risk analyses, detection, and control, all designed with inherent cyber awareness. The system’s architecture extends from end-devices in the field through to control center applications, establishing a secure and resilient control framework that addresses the challenges posed by diverse devices of unknown trustworthiness connecting to modern power systems. Through this innovative approach to deep cyber-physical situational awareness, the CYPRES project not only advances the state-of-the-art in energy infrastructure protection but also establishes a new paradigm for how EMS can be designed, deployed, and operated in an increasingly complex threat landscape. The findings and developments from this project provide crucial insights for stakeholders across the energy sector, offering a blueprint for enhancing the reliability and resilience of our nation’s critical energy infrastructure in the face of evolving cyber threats.

24 POWER TRANSMISSION AND DISTRIBUTION

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)

LLM Generation of Online Courses from a Curated Set of Documents in the Nuclear Safeguards Domain

A multidisciplinary team at Argonne National Laboratory explores the application of advanced technologies to enhance knowledge transfer and retention within the nuclear safeguards domain. Specifically, it examines the feasibility of leveraging secure large language models (LLMs) to streamline the creation of e-learning modules for the U.S. National Nuclear Security Administration (NNSA) Office of International Nuclear Safeguards (NA-241). The initiative addresses the critical need for preserving institutional memory and accelerating skill development amidst the imminent retirement of senior professionals in the field in addition to supporting good knowledge management practices. The project integrates instructional design theory with cutting-edge AI technologies to transform curated document sets from the Safeguards Knowledge Repository (SKR) into modular online courses. By automating the generation of learning objectives and instructional content, the effort aims to reduce manual effort while maintaining high-quality educational outcomes. A limited measure of human supervision, however, ensures accuracy, relevance, and alignment with NNSA’s strategic priorities. Key findings highlight the potential of AI-assisted course generation to support safeguards professionals by creating structured, interactive learning experiences. The report underscores the importance of SME validation to address limitations in AI-generated content, such as terminology errors and gaps in coverage. Recommendations include adopting a structured workflow combining LLM acceleration with expert oversight to ensure accuracy, usability, and alignment with learner needs. This work demonstrates Argonne’s commitment to advancing national security and scientific excellence through innovative knowledge management solutions.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

BuildingSync® v.2.7.0 (released 9.11.2025) [SWR-18-28]

BuildingSync® is a building data exchange schema to better enable integration between software tools and building data workflows. The schema's original use case was focused on commercial building energy audits; however, several additional use cases have been realized including building energy modeling and more high-level generic building data exchange. Version 2.7.0 adds new elements for file attachment feature and FederalBuilding, and generalizes usage of Optional Elements (e.g. EquipmentCondition, EquipmentID) to all assets/systems. BuildingSync helps streamline the data exchange process, improving the value of the data, minimizing duplication of effort for subsequent building data collection efforts (including audits), and facilitating the achievement of greater energy efficiency. This in done in part by standardizing on (a) reporting audits in an electronic format, (b) tracking proposed, implemented, and discarded energy conservation measures, and (c) storing building characteristics (at multiple levels) for audits, benchmarking, and building energy analysis. BuildingSync has several documents and tools available to help users understand how to best leverage BuildingSync. The list below are only a subset of the resources available. If new resources are discovered, then feel free to create a new pull request with the additions. Generic BuildingSync information is available on the DOE website and the project website. BuildingSync Examples - These examples are kept up to date and show a wide range of implementations. Any new update to BuildingSync is required to pass validation on these example files. BuildingSync Use Case Validator allows for users to determine if their instance complies with a specific use case for BuildingSync by checking if the required elements are implemented in an uploaded instance. An API is also provided for automated integration into other tools. Also, the website contains an easy way to view the entirety of the schema and how elements relate to the Building Exchange Data Exchange Specification. The Validator is open sourced here Use Case TestSuite provides a Python package for easier generation of BuildingSync use cases. BuildingSync use cases depend on the generation of schematron documents, which is time-consuming and difficult to implement well. The TestSuite allows users to define a use case using a more palatable CSV template, which it then turns into a Schematron document. The source code is available here. BuildingSync to OpenStudio/EnergyPlus. The translator is open sourced here. This project will translate a Level 1 (and partial Level 2) ASHRAE Energy Audit to a fully defined OpenStudio and EnergyPlus model. This project is in early Beta testing and any feedback is welcome!

Long, Nicholas [National Renewable Energy Lab. (NR

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES

CONFLUX: A standardized framework to calculate reactor antineutrino flux

Nuclear fission reactors are abundant sources of antineutrinos for neutrino physics experiments. The flux and spectrum of antineutrinos emitted by a reactor can indicate its activity and composition, suggesting potential applications of neutrino measurements beyond fundamental scientific studies that may be valuable to society. The utility of reactor antineutrinos for applications and fundamental science is dependent on the availability of precise predictions of these emissions. For example, in the last decade, disagreements between reactor antineutrino measurements and models have inspired revision of reactor antineutrino calculations and standard nuclear databases as well as searches for new fundamental particles not predicted by the Standard Model of particle physics. Past predictions and descriptions of the methods used to generate them are documented to varying degrees in the literature, with different modeling teams incorporating a range of methods, input data, and assumptions. The resulting difficulty in accessing or reproducing past models and reconciling results from differing approaches complicates the future study and application of reactor antineutrinos. The CONFLUX (Calculation Of Neutrino FLUX) software framework is a neutrino prediction tool built with the goal of simplifying, standardizing, and democratizing the process of reactor antineutrino flux calculations. CONFLUX includes three primary methods for calculating the antineutrino emissions of nuclear reactors or individual beta decays that incorporate common nuclear data and beta decay theory. The software is prepackaged with the current nuclear databases, including ENDF.B/VIII, JEFF-3.3, and ENSDF, and it includes the capability to predict time-dependent reactor emissions, adjust nuclear database or beta decay inputs/assumptions, and propagate related sources of uncertainty. Here, this paper describes the CONFLUX software structure, details the methods used for flux and spectrum calculations, and provides examples of potential use cases.

Zhang, Xianyi [Lawrence Livermore National Laborat