Search NASA⌕ Search

SEARCH · Search NASA

Results for “Document Generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

Gold-Standard Chemical Database 137 (GSCDB137): A Diverse Set of Accurate Energy Differences for Assessing and Developing Density Functionals

We present GSCDB137, a rigorously curated benchmark library of 137 data sets (8377 entries) covering main-group and transition-metal reaction energies and barrier heights, (intra- and intermolecular) noncovalent interactions, dipole moments, polarizabilities, electric-field response energies, and vibrational frequencies. Legacy data from GMTKN55 and MGCDB84 have been updated to today's best reference values; redundant or low-quality points were removed, and many new, property-focused sets were added. Testing 29 popular density functional approximations (DFAs) confirms the expected Jacob's-ladder hierarchy overall but also reveals notable exceptions: functional performance for frequencies and electric-field properties correlates poorly with that for other ground-state energetics. ωB97M-V and ωB97X-V are the most balanced hybrid meta-GGA and hybrid GGA, respectively; B97M-V and revPBE-D4 lead the meta-GGA and GGA classes. Double hybrids lower mean errors by about 30% versus their hybrid analogues but demand careful frozen-core, basis set, and spin contamination treatment. GSCDB137 offers a comprehensive, openly documented platform for rigorous validation of DFA and universal machine learning potentials, and training of the next generation of exchange-correlation functionals.

Liang, Jiashu [University of California, Berkeley,↗

DE‐SC0022708 RESEARCH PERFORMANCE FINAL REPORT

GismoPower’s Final Research Report documents the outcomes of a DOE SBIR Phase II project focused on advancing the MEGA® (Mobile Electricity Generating Appliance), a trailerable, plug-in solar canopy appliance designed to deliver appliance-class electricity generation for homes, small businesses and renters. The project’s core objective was to remove the technical and regulatory barriers that have historically prevented plug-in solar systems from being safely certified, permitted, and interconnected in the United States.

14 SOLAR ENERGY↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

LUCID Thrust 1 - Dataset Identification and Biodata Catalog Creation

The LUCID DOE consortium, part of the Department of Energy’s Biological and Environmental Research (BER) program, advances Low Dose Radiation (LDR) research through multidisciplinary efforts across seven key thrusts. This document focuses on Thrust 1, which centers on the creation of curated multimodal population health datasets and supports broader efforts within the LUCID program, including AI-based hypothesis generation, experimental design, and the study of LDR-induced health risks. Specifically, it describes the identification and cataloging of Thrust 1’s curated LDR datasets and biodata, emphasizing their critical role in supporting various research thrusts within the consortium, with potential applications in healthcare and public policy. In addition, the document includes an evaluation of three Large Language Models (LLMs)—GPT-4, SOLAR-10B, and Mixtral-8x7B—based on their ability to extract features from 25 LDR studies. The results indicate that GPT-4 performed the best, while Mixtral-8x7B demonstrated limited knowledge. Overall, this work advances understanding in radiation protection, risk assessment, and medical treatments, while providing valuable resources for researchers, educators, and policymakers.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

OC7 Phase I Definition Document

The Offshore Code Comparison Collaboration 7 (OC7) project is organized under the International Energy Agency Wind Technology Collaboration Programme Task 56 with an objective to evaluate and enhance the predictive accuracy of engineering-level modeling tools used in the design of offshore wind energy systems. Phase I of OC7 is focused on improving the models and modeling practices of hydrodynamic viscous loads on floating offshore wind turbine platforms. In alignment with this goal, Work Package 1.1 of OC7 Phase I is formulated to investigate the modeling of hydrodynamic viscous loads on several different geometric components commonly encountered with floating offshore wind platform designs, including cylindrical columns, heave plates, and rectangular pontoons. Work Package 1.1 also explores the dependence of hydrodynamic coefficients on the sea state to drive toward practical guidance on how these coefficients can be selected or adjusted for different conditions. This report outlines the motivation and objectives behind each subphase of OC7 Phase I, along with the necessary technical specifications and load case definitions to guide the project participants. It also serves as an important part of the project documentation for future modelers who would like to reproduce this work or make use of the data and information generated from the OC7 project.

17 WIND ENERGY↗

DMTN-277: The Monster: A reference catalog with synthetic ugrizy-band fluxes for the Vera C. Rubin observatory

In order to facilitate bootstrap photometric calibrations of early Rubin Observatory data we have created an all sky reference catalog called The Monster. This reference catalog uses a rank-ordered set of other reference catalogs to generate synthetic ugrizy-band fluxes that can be used calibrate images processed with the LSST science pipelines. This document describes the methodology used to create The Monster, documents the input external reference catalogs, and performs basic data validation of the first version of The Monster.

79 ASTRONOMY AND ASTROPHYSICS↗

Tank 11H Low Temperature Aluminum Dissolution and Inhalation Dose Potential Analyses at Savannah River Site – 26018

Currently, there is approximately 34 million gallons of high-level radioactive tank waste in the Tank Farm at the Savannah River Site (SRS). The ultimate goal of operations at the Tank Farm is to remove the high level waste (HLW) from the tanks followed by stabilization of the waste through vitrification of the HLW into glass or grouting the decontaminated waste into saltstone. After bulk removal of the HLW consisting of sludge, saltcake, and supernatant, further efforts are made to reduce the residual waste present in the tank in order to declare preliminary cease waste removal (PCWR) signifying completion of HLW removal. These reduction efforts can include tank washing to remove soluble salts and radioisotopes and dissolution of solids including aluminum. Aluminum in the form of gibbsite and boehmite is relatively insoluble in water. Through addition of aqueous sodium hydroxide, the aluminum can be dissolved at mild temperatures. In order for the waste tank to meet closure mode requirements of the Concentration, Storage, and Transfer Facilities (CSTF), which includes the Tank Farm, Documented Safety Analysis (DSA), a component of the safety basis, the inhalation dose potential (IDP) and the radiolytic hydrogen generation rate of the stored waste must be demonstrated to be lower than their respective designated limits. These parameters are calculated from measured radiochemical analyses of isotopes that emit a high amount of radioactivity including Cs-137, Sr-90, Pu-238, Pu-239, Pu-240, Pu-241, Am-241, and Cm-244. Following the low temperature aluminum dissolution (LTAD) process, Tank 11H slurry samples were pulled from the tank and sent to Savannah River National Laboratory (SRNL) to measure the extent of aluminum dissolution, hydroxide concentration, densities of slurry and supernatant, weight percent solids analyses, and radionuclide activities. The analyses of the composite sample found that approximately 90% of the total aluminum in the slurry was dissolved, indicating successful reduction of the insoluble aluminum in the waste tank. Additionally, the weight percent insoluble solids (slurry basis) measurement of the composite sample was found to be approximately 1%, demonstrating that minimal solids still remain in the tank. Finally, the radiochemical analyses of the composite sample determined that the waste contents of the tank met the IDP and radiolytic hydrogen generation rate requirements of the CSTF DSA. These measurements have shown that the LTAD process in Tank 11H was successful in waste reduction efforts and a positive step towards declaring PCWR and tank closure at SRS.

Dekarske, John [Savannah River National Laboratory↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ERF: Energy Research and Forecasting Model

High performance computing (HPC) architectures have undergone rapid development in recent years. As a result, established software suites face an ever increasing challenge to remain performant on and portable across modern systems. Many of the widely adopted atmospheric modeling codes cannot fully (or in some cases, at all) leverage the acceleration provided by General-Purpose Graphics Processing Units, leaving users of those codes constrained to increasingly limited HPC resources. Energy Research and Forecasting (ERF) is a regional atmospheric modeling code that leverages the latest HPC architectures, whether composed of only Central Processing Units (CPUs) or incorporating GPUs. ERF contains many of the standard discretizations and basic features needed to model general atmospheric dynamics. The modular design of ERF provides a flexible platform for exploring different physics parameterizations and numerical strategies. ERF is built on a state-of-the-art, well-supported, software framework (AMReX) that provides a performance portable interface and ensures ERF's long-term sustainability on next generation computing systems. This paper details the numerical methodology of ERF, presents results for a series of verification/validation cases, and documents ERF's performance on current HPC systems. The roughly 5× speed up of ERF (using GPUs) over Weather Research and Forecasting (CPUs only) for a 3D squall line test case highlights the significance of leveraging GPU acceleration.

17 WIND ENERGY↗

Non-Refrigerated Warehouse Design Guide [Slides]

Warehouses typically have a low energy use relative to their floor area. This make warehouses ideal to install solar panels and offset on-site energy use and have a surplus of energy if roof area is maximized. This surplus can be used to provide things like grid services or charge electric vehicle fleets. This document is intended to provide owners and operators of warehouse information on designing efficient/decarbonized warehouses and maximize on-site generation. The information here references outside publications to provide specific guidance, such as the ASHRAE Advanced Energy Design Guides, but tailors the information specifically to warehouses. This guide provides owners a better understanding of warehouse design as well as more specific information for themselves or outside designers to reference.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Infrastructure-Based Cooperative Perception at a Traffic Intersection: Overview and Challenges

Traffic intersections are crucial and challenging nodes in transportation networks where multiple lanes of vehicles and pedestrians converge. About one-quarter of traffic fatalities and about one-half of all traffic injuries in the United States happen at traffic intersections . Effective management of these intersections is important to ensure safety and efficiency of all users - vehicles, pedestrians, cyclists, and vulnerable road users (VRUs). With advancements in sensor perception technologies such as radar, light detection and ranging (lidar), and cameras, traffic intersections are developing into dynamic and data-rich environments. By using these data to create a real-time digital twin, we can enable real-time data-driven decision making and a range of applications such as sharing perception information to connected vehicles (CVs) and connected autonomous vehicles (CAVs), safety affirmative signaling, and curb optimizing to improve efficiency and enhance safety.This paper presents an overview of the concept and examines the challenges involved in implementing an infrastructure-based cooperative perception engine at a traffic intersection. In addition to outlining the physical components, this study also addresses important challenges involved in a multi-sensor system. We present results from deploying the National Renewable Energy Laboratory's (NREL's) Infrastructure Perception and Control (IPC) mobile trailer at a traffic intersection in the city of Colorado Springs, Colorado, USA that employed multiple radars and lidars to capture the data. This study provides necessary practical learning for the Cooperative Driving Automation (CDA) and traffic engineering communities for next-generation infrastructure-based cooperative perception that promises improvements in signal control for optimized traffic flow, among other applications, and documents findings for ongoing research and development efforts in other areas.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

Updated General Aviation Non-Airport Crash Density Values Using Data Obtained from the U.S. National Transportation Safety Board for 2000 through 2019

This report documents the calculation of the General Aviation (GA) non-airport crash density for Department of Energy (DOE) sites and power generating facilities licensed by the Nuclear Regulatory Commission (NRC) using the methodology outlined in DOE-STD-3014-2006, Accident Analysis for Aircraft Crash Into Hazardous Facilities. The results of this report will be used by the American Nuclear Society (ANS) for the development of ANS standard 2.36, Accident Analysis for Aircraft Crash into Reactor and Nonreactor Nuclear Facilities.

99 GENERAL AND MISCELLANEOUS↗

Integrating Energy-Efficient Computing with Computational Research to Accelerate Energy Technology

NREL's computational sciences center hosts the largest high performance computing (HPC) capabilities dedicated to energy research while functioning as a living laboratory for energy-efficient computing. NREL's HPC capabilities support the research needs of the Department of Energy's Office of Energy Efficiency and Renewable Energy (EERE). In ten years of operation, HPC use in EERE-sponsored research has grown by a factor of 30, including work in electricity generation, energy efficiency, transportation, and energy system modeling. This paper analyzes this research portfolio, providing examples of individual use cases. The paper documents NREL's history of operating one of the world's most energy-efficient data centers while examining pathways to reduce economic and environmental impact beyond reduction of Power Usage Efficiency (PUE). This paper concludes by examining the unique opportunities created for accelerating improvements in data center efficiency created by combining an HPC system dedicated to energy research and a research program in energy-efficient computing.

97 MATHEMATICS AND COMPUTING↗

ESAC v1: Enhanced AI-Powered Chatbot for EQ-SANS Experiment Automation Improvements and Updates

ESAC (EQ-SANS Assisting Chatbot) is an advanced AI-powered application designed to streamline the workflow of neutron scattering experiments at the Spallation Neutron Source (SNS). This report documents the significant advancements in ESAC v1, which include the integration of a combined In-Context Learning (ICL) + Retrieval-Augmented Generation (RAG) capability, a robust integrated development environment, and standalone executable distribution. These enhancements address the limitations of the original version, making ESAC v1 a transformative tool for researchers. The report also discusses the technical challenges encountered during development and their resolution, highlighting the impact of these improvements on the neutron scattering research community.

36 MATERIALS SCIENCE↗

2024 Annual Technology Baseline (ATB) Cost and Performance Data for Electricity Generation Technologies

These data provide the 2024 update of the Electricity Annual Technology Baseline (ATB). Starting in 2015 NREL has presented the ATB, consisting of detailed cost and performance data, both current and projected, for electricity generation and storage technologies. The ATB products now include data (Excel workbook, Tableau workbooks, and structured summary csv files), as well as documentation and user engagement via a website, presentation, and webinar. Starting in 2021, the data are cloud optimized and provided in the OEDI data lake. The data for 2015 - 2020 are can be found on the NREL Data Search Page. The website documentation can be found on the ATB Website.

Array↗

Common practices for quantifying methane emissions from plumes detected by remote sensing

This document provides a set of community-accepted practices for quantifying methane emissions based on plumes detected via spectroscopic remote sensing. Its primary goal is to promote consistency in the generation, validation, reporting, and quality assessment of methane emission estimates derived from remote sensing radiances. Developed by subject matter experts with deep experience across all stages of the measurement process, this guidance reflects a critical evaluation of current methodologies and highlights key practices needed to produce reliable, interoperable, and traceable products. The focus is specifically on methane emissions quantified from distinct plumes originating from localized sources, rather than diffuse emissions spread over large regions, which are beyond the scope of this work. This document is intended to serve both data producers and users. For producers, it offers a framework for aligning with field-recognized standards to ensure their outputs meet rigorous quality and transparency criteria. For users, it provides a reference to assess dataset fitness-for-purpose by highlighting essential metadata, assumptions, and methodological choices that underpin emission estimates. By fostering a shared understanding of best practices, this work aims to enhance comparability, confidence, and utility of remotely sensed methane emission products.

54 ENVIRONMENTAL SCIENCES↗