Search NASA⌕ Search

SEARCH · Search NASA

Results for “document repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Developing Concepts of Operations Using Multi-Step Tool Techniques With Large Language Models

The National Aeronautics and Space Administration (NASA) Air Mobility Pathfinders (AMP) project is developing and evaluating concepts of operations (ConOps) for safe, secure, and scalable Urban Air Mobility (UAM) operations. The AMP project’s Operational Concepts, Architecture, and Requirements Integration (OCARI) Team is using a Model Based System Engineering (MBSE) approach for integration, interoperability, and traceability of Advanced Air Mobility (AAM) ecosystems centered around urban air taxi services. The team’s goal is to define structures and behaviors needed for system feasibility, readiness, and interoperability, establish a UAM knowledge base, and trace and validate assumptions and requirements relevant to AAM. NASA Langley Research Center (LaRC) is spearheading an innovative digital engineering approach to integrate, communicate, and facilitate the research of multi-modal transportation systems. The Knowledge-based Digital Platform (KbDP) is a concept being developed that ties the workflows of Project Managers (PM), Principal Investigators (PI), and System Engineers together across organizational boundaries. It does so through the management of an information database defined by mathematical, data science, and system engineering principles. Machine Learning (ML) algorithms play a key role in this concept by extracting meaningful knowledge from relational and graph databases, document repositories, and system artifacts, which the human user leverages to greatly improve the efficiency and effectiveness of their research. Recent advancements in the field of Large Language Models (LLMs), specifically models trained for tool use, such as Command-R , now allow for the reliable implementation of single-step and multi-step tool-centric systems. These techniques provide the LLM with a set of tools, in our case Python functions, that can be called on to answer a much wider range of questions compared to LLMs implemented using a traditional single-source or Retrieval Augmented Generation (RAG) approach. Through this method, the LLM can pull information from multiple data sources, such as relational or graph databases, document repositories, application programming interfaces (APIs), and SysML artifacts depending on the user’s question. The LLM can also output the information in a variety of different formats, using output generation tools, such as CSV, UML, or SysML artifacts. Additionally, tools can be assigned roles and can work together to provide answers to queries in an “agent” like approach, similar to that implemented by Microsoft’s AutoGen framework where different agents can converse with each other to accomplish tasks. Previously, our team developed a chatbot system with “agent like” functionality in the form of different “modes” the user could select from a user interface (UI), this architecture can be seen on the left in figure 1. Three different modes were implemented, the first mode allowed the LLM to utilize the structures and algorithms within a graph database to trace UAM requirements. The second mode gave the LLM access to a vector search capable of providing relevant information from thousands of document pages related to UAM ConOps and requirements. The third mode served as a general assistant where users could enter open-ended questions and custom prompts to utilize the LLM for different use-cases. This system improved the process surrounding generating and analyzing information related to UAM requirements, however, the implementation provided a clunky user experience. Users were required to know what mode to select within the UI in advance before entering their question to the selected tool. Moreover, the different tools were isolated from each other, they lacked bidirectional links that would allow for tools to collaborate to generate better responses. Our team is working on a new architecture, seen on the right in the below figure, with the goal to address many of the UX shortcomings of our original system while improving the accuracy and depth of responses from the LLM. This new system will automatically select the appropriate tool to use based off the user’s question. Each tool will be capable of calling on any of the other tools available to the LLM, resulting in a collaborative pipeline where tools can pass data between other tools until enough data is received to generate an answer to the user’s question. Using a locally deployed, open-source, LLM, the NASA OCARI team, in collaboration with Collins Aerospace, will implement a prototype application that will bridge knowledge across multiple sources to assist System Engineers (SEs) with requirements discovery and tracing, research question and use case identification, and assumption validation. Such a system will also allow SEs to more easily, and intuitively, explore the AAM ecosystem, ultimately improving the efficiency and effectiveness of the SE's research and decision-making processes surrounding ConOps development and validation. In this session, our team will provide a video demonstration of our new prototype architecture in action. We will also present an overview of our prototype system architecture and talk about its advantages over traditional LLM deployments along with how those advantages can provide additional value to the field of System Engineering.

systems engineering↗

Final Technical Report

Statement of the problem or situation that is being addressed in your application. The DOE and its national laboratories developed the Home Energy Score™ (HES) to encourage homeowners to improve their energy performance, lower costs and to share energy information through the MLS listing, appraisal, and financing channels. While the HES is an instrumental tool, it is currently underutilized and consists of technical, structural and sector barriers which need to be addressed in order to scale and many energy efficiency contractors are understandably overwhelmed by the added time and effort and lack of incentive to sell and deliver deep retrofit projects while simultaneously meeting the DOE HES program requirements; consequently, contractors may decide to forgo participation. Home Energy Rating System (HERS) Raters have the opportunity to play the critical Assessor role in producing a Home Energy Score (HES); this role has immense potential but currently is unfulfilled. Lastly, while utilities are interested in their customer base achieving greater energy efficiency, especially to help offset growing residential loads in states like California that are accelerating electrification, utilities do not have access to the market actors who are on the front line of influence to homeowners or review and approve their permits: HERS Raters, assessors, contractors and building departments. General statement of how this problem is being addressed: ConSol will integrate the Home Energy Score™ (HES) to its State of California, approved home energy rating services (HERS) platform (CHEERS) to develop a single tool for contractors nationwide to assess, record and install recommended cost, energy, and emissions saving measures to the 140 million single-family homes throughout the U.S. and 14 million homes in California (CHEERS+HES). The CHEERS high fidelity energy code permitting data will be integrated with HES for simple, accurate, easy-to-use home energy estimation and analysis and will directly gain access to the retrofit and renovations markets with the same upgraded platform. This innovative project will assist the utilities in supporting existing homes in their jurisdictions with HES and develop measures to improve energy efficiency and reduce emissions. How is this problem being addressed? What is the overall project approach? In effort to expand the Home Energy Score™ (HES) by increasing the use of aggregable home energy asset data, ConSol proposes to integrate the DOE HES via Application Programming Interface (API) to its State of California approved home energy rating services platform (CHEERS). Once the CHEERS platform and HES are integrated (CHEERS+HES), this enhanced platform will be instantly available and actively deployed via Phase 1 pilot to HERS Raters, assessors and contractors in California to market-test the solution, understand the rate of adoption and identify opportunities for improvement prior to scaling nationally. The CHEERS high fidelity energy code permitting data will be integrated with HES for simple, accurate, easy-to-use home energy estimation and analysis and will directly gain access to the retrofit and renovations markets with the same upgraded platform. This innovative project will assist the building industry and homeowners with an easy-to-use assessment if energy and carbon impacts of existing homes, and assist the utilities in supporting existing homes in their jurisdictions with HES to improve energy efficiency and reduce emissions. What is to be done in Phase I? During Phase I of this proposed project, ConSol will (1) design software architecture that links CHEERS to the Home Energy ScoreTM via API, (2) solicit partnership from one or more California utilities for a regional pilot, (3) test the new software with its HERS Raters and contractor network in the partnership utility jurisdiction, (4) launch a pilot version of the newly developed software with HERS Raters and contractors in the utility territory, and (5) explore California’s GoGreen energy efficiency homeowner lending program in parallel with the pilot. Commercial Applications and Other Benefits. Summarize the future applications or public benefits if the project is carried over into Phase II or Phase III and beyond. The CHEERS+HES commercialized product will be ready for national market scale following a successful Phase 1 performance. The CHEERS+HES adoption is estimated to reach a 5% adoption growth rate versus the 110,000 baseline, starting in Year 1 after Phase I completion, and continuing each year. As a direct benefit to the DOE, CHEERS will set a goal of 100,000 Home Energy Score assessments for existing home alterations within the first 10 years following Phase 1 performance. The technical benefits of this proposed project include the harmonized, automated, and seamless integration of the DOE HES into the widely used and market leading California energy registry, CHEERS. The social benefits include the aggregate energy, cost and GHG savings by allowing the broader public streamlined access to the CHEERS+HES measurement and the energy efficiency recommended measures that may result. Key Words: Home Energy ScoreTM (HES); Application Programming Interface (API); Home Energy Rating Services (HERS); HERS Raters; contractors; assessors; existing homes, energy asset data; cost, energy, and emissions saving measures; energy code (Title 24) compliance; document repository; utilities; pilot; newly developed software; energy efficiency; homeowner. Summary for Members of Congress: The DOE Home Energy Score™ (HES) is a tool to encourage homeowners to improve their energy performance, lower costs and share energy information but is underutilized and consists of barriers which need to be addressed in order to scale. In effort to expand the HES, CHEERS, Inc. will integrate the HES to its State of California, approved home energy rating services (HERS) platform (CHEERS) to develop a single tool for contractors nationwide to assess, record and install recommended cost, energy, and emissions saving measures to the 140 million single-family homes throughout the U.S. and 14 million homes in California.

Application Programming Interface (API)↗

InvestigationOrganizer: The Development and Testing of a Web-based Tool to Support Mishap Investigations

InvestigationOrganizer (IO) is a collaborative web-based system designed to support the conduct of mishap investigations. IO provides a common repository for a wide range of mishap related information, and allows investigators to make explicit, shared, and meaningful links between evidence, causal models, findings and recommendations. It integrates the functionality of a database, a common document repository, a semantic knowledge network, a rule-based inference engine, and causal modeling and visualization. Thus far, IO has been used to support four mishap investigations within NASA, ranging from a small property damage case to the loss of the Space Shuttle Columbia. This paper describes how the functionality of IO supports mishap investigations and the lessons learned from the experience of supporting two of the NASA mishap investigations: the Columbia Accident Investigation and the CONTOUR Loss Investigation.

Carvalho, Robert F.↗

Automatic MCNP File Generation for Cherenkov Imaging Simulations [Slides]

Overview: Creating physics models to accurately simulate a variety of spent fuel assemblies. These characterized simulations will then be used to create a well documented repository that encompasses a wide variety of different fuel assemblies, burn-up and spent fuel pool conditions, and defects. The final project will use our simulations to train an AI, and for it to be effective it has to have a lot of reliable data; Automating an arbitrary amount of created spent fuel pins with materials, that will be passed into AI models for it to "learn" what a good spent nuclear fuel cell looks like and what a non-spent nuclear fuel cell looks like.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Accident/Mishap Investigation System

InvestigationOrganizer (IO) is a Web-based collaborative information system that integrates the generic functionality of a database, a document repository, a semantic hypermedia browser, and a rule-based inference system with specialized modeling and visualization functionality to support accident/mishap investigation teams. This accessible, online structure is designed to support investigators by allowing them to make explicit, shared, and meaningful links among evidence, causal models, findings, and recommendations.

Keller, Richard↗

Concept document of the repository-based software engineering program: A constructive appraisal

A constructive appraisal of the Concept Document of the Repository-Based Software Engineering Program is provided. The Concept Document is designed to provide an overview of the Repository-Based Software Engineering (RBSE) Program. The Document should be brief and provide the context for reading subsequent requirements and product specifications. That is, all requirements to be developed should be traceable to the Concept Document. Applied Expertise's analysis of the Document was directed toward assuring that: (1) the Executive Summary provides a clear, concise, and comprehensive overview of the Concept (rewrite as necessary); (2) the sections of the Document make best use of the NASA 'Data Item Description' for concept documents; (3) the information contained in the Document provides a foundation for subsequent requirements; and (4) the document adequately: identifies the problem being addressed; articulates RBSE's specific role; specifies the unique aspects of the program; and identifies the nature and extent of the program's users.

Source record↗

Repository-based software engineering program: Concept document

This document provides the context for Repository-Based Software Engineering's (RBSE's) evolving functional and operational product requirements, and it is the parent document for development of detailed technical and management plans. When furnished, requirements documents will serve as the governing RBSE product specification. The RBSE Program Management Plan will define resources, schedules, and technical and organizational approaches to fulfilling the goals and objectives of this concept. The purpose of this document is to provide a concise overview of RBSE, describe the rationale for the RBSE Program, and define a clear, common vision for RBSE team members and customers. The document also provides the foundation for developing RBSE user and system requirements and a corresponding Program Management Plan. The concept is used to express the program mission to RBSE users and managers and to provide an exhibit for community review.

Source record↗

MSD CoP Webinar: Using Meta-Repositories To Facilitate Open Science in MSD Research

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Abstract: In this webinar, we will describe the use of a GitHub meta-repository, for documenting and disseminating the tools and data supporting MSD publications. The goal is to make it easier for others to understand the flow of data and code through your experiment and to be able to reproduce your results and figures with only the information you have provided for them. The webinar will cover the role of open science in the MSD community, the origins and purpose of meta-repositories, and step-by-step instructions and best practices for building a meta-repository starting from the GitHub template (https://github.com/IMMM-SFA/metarepo). We will also discuss how to leverage MSD-LIVE (https://msdlive.org/) in your meta-repository, provide links to numerous examples you can learn from, and discuss the role of meta-repositories in the IM3 project's open science mandates. Presenters : Chris R. Vernon, Casey D. Burleyson, Jennie Rice, and Mengqi Zhao Moderator: Pat M. Reed (MSD CoP Facilitation Team) This webinar was held on: February 22nd, 2024 from 2-3 PM ET

Open Science↗

NELS 2.0 - A general system for enterprise wide information management

NELS, the NASA Electronic Library System, is an information management tool for creating distributed repositories of documents, drawings, and code for use and reuse by the aerospace community. The NELS retrieval engine can load metadata and source files of full text objects, perform natural language queries to retrieve ranked objects, and create links to connect user interfaces. For flexibility, the NELS architecture has layered interfaces between the application program and the stored library information. The session manager provides the interface functions for development of NELS applications. The data manager is an interface between session manager and the structured data system. The center of the structured data system is the Wide Area Information Server. This system architecture provides access to information across heterogeneous platforms in a distributed environment. There are presently three user interfaces that connect to the NELS engine; an X-Windows interface, and ASCII interface and the Spatial Data Management System. This paper describes the design and operation of NELS as an information management tool and repository.

Smith, Stephanie L.↗

Space Telecommunications Radio System (STRS) Application Repository Design and Analysis

The Space Telecommunications Radio System (STRS) Application Repository Design and Analysis document describes the STRS application repository for software-defined radio (SDR) applications intended to be compliant to the STRS Architecture Standard. The document provides information about the submission of artifacts to the STRS application repository, to provide information to the potential users of that information, and for the systems engineer to understand the requirements, concepts, and approach to the STRS application repository. The STRS application repository is intended to capture knowledge, documents, and other artifacts for each waveform application or other application outside of its project so that when the project ends, the knowledge is retained. The document describes the transmission of technology from mission to mission capturing lessons learned that are used for continuous improvement across projects and supporting NASA Procedural Requirements (NPRs) for performing software engineering projects and NASAs release process.

information retrieval↗

Hermes-3: Multi-component plasma simulations with BOUT++

A new open source tool for fluid simulation of multi-component plasmas is presented, based on a flexible software design that is applicable to scientific simulations in a wide range of fields. Hermes-3 is built on plasma simulation framework BOUT++, consolidating earlier SD1D and Hermes models into a single code that can be configured at run-time to solve plasma models in 1D, 2D or 3D, either for transport (steady-state) or turbulent (time-evolving) problems, with an arbitrary number of ion and neutral species. Here, we describe the improved numerical algorithms and software design that have been implemented in Hermes-3. To demonstrate the capabilities of this tool, applications relevant to the boundary of tokamak plasmas are presented: 1D simulations of diveror plasmas evolving equations for all charge states of neon and deuterium; 2D transport simulations of tokamak equilibria in single-null X-point geometry with plasma ion and neutral atom species; and simulations of the time-dependent propagation of plasma filaments (blobs). Hermes-3 is publicly available on Github under the GPL-3 open source license. The repository includes documentation and a suite of unit, integrated and convergence tests.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

FrESCO: Framework for Exploring Scalable Computational Oncology

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information from electronic health records. This data is then used to create summary statistics about cancer incidence and mortality to facilitate population health monitoring. Extracting phenotypic information from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating the information extraction process from cancer pathology reports has the potential to improve data quality by extracting information in a consistent manner across registries. It can also improve patient outcomes by reducing the time from diagnosis, enabling rapid case ascertainment for clinical trials. Here we present FrESCO, a modular deep-learning natural language processing (NLP) library initially designed for extracting pathology information from clinical text documents. This repository is not solely limited to clinical medical text, but may also be used by researchers just getting started with NLP methods and those looking for a robust solution for their classification problems.

60 APPLIED LIFE SCIENCES↗

Macrofossil extinction patterns at Bay of Biscay Cretaceous-Tertiary boundary sections

Researchers examined several K-T boundary cores at Deep Sea Drilling Project (DSDP) core repositories to document biostratigraphic ranges of inoceramid shell fragments and prisms. As in land-based sections, prisms in the deep sea cores disappear well before the K-T boundary. Ammonites show a very different extinction pattern than do the inoceramids. A minimum of seven ammonite species have been collected from the last meter of Cretaceous strata in the Bay of Biscay basin. In three of the sections there is no marked drop in either species numbers or abundance prior to the K-T boundary Cretaceous strata; at the Zumaya section, however, both species richness and abundance drop in the last 20 m of the Cretaceous, with only a single ammonite specimen recovered to date from the uppermost 12 m of Cretaceous strata in this section. Researchers conclude that inoceramid bivalves and ammonites showed two different times and patterns of extinction, at least in the Bay of Biscay region. The inoceramids disappeared gradually during the Early Maestrichtian, and survived only into the earliest Late Maestrichtian. Ammonites, on the other hand, maintained relatively high species richness throughout the Maestrichtian, and then disappeared suddenly, either coincident with, or immediately before the microfossil extinction event marking the very end of the Cretaceous.

Ward, Peter D.↗

REFSafE: A RAG-Enabled Framework for Predictive Risk Analysis and Automated Safety Report Generation in Mission-Critical Environments

Operational safety in mission-critical environments requires AI systems that are accurate, interpretable, and resistant to hallucination. We present an agentic Retrieval-Augmented Generation (RAG) framework, REFSafe, for grounded hazard analysis and automated safety report generation. The system integrates Large Language Models (LLMs) with structured operational data, historical incident repositories, policy documents, and external authoritative sources. Through iterative agentic reasoning, the framework retrieves, verifies, and synthesizes evidence prior to generation, enforcing citation-backed outputs with explicit source attribution (documents, links, and prior events) to ensure traceability and trust. To mitigate hallucinations and unsupported claims, all risk assessments and forecasts are constrained to retrieved evidence, with confidence signals derived from retrieval relevance and source consistency. A transparent pipeline enables subject matter experts (SMEs) to validate predictions, and provide structured feedback, forming a continuous performance calibration loop. Preliminary deployment demonstrates improved reliability in hazard detection and safety/vulnerability report generation. This work advances trustworthy, evidence-grounded AI for predictive safety intelligence in mission-critical operations.

Das, Sanjay [ORNL] (ORCID:0009000542591915)↗

Natural Language Processing Techniques for Intelligent Knowledge Management of Safety Reports

Safety, failure, and incident reports are common artifacts across various domains, including aviation and wildfire response. These reports are often mandatory to submit, resulting in the culmination of large repositories of text-based documents. Simultaneously, these reports and corresponding repositories are often only manually analyzed and queried by users via out-of-date search engines. As a consequence, we have been developing the Manager for Intelligent Knowledge Access (MIKA) toolkit, which uses natural language processing to improve information access and reuse. In this presentation, we discuss natural language processing techniques for knowledge discovery and apply these methods to a repository of aerial wildfire mishap reports. Two methods are used for knowledge discovery: topic modeling and named-entity recognition. We use topic modeling to identify hazards and perform a trend analysis to produce a data-driven risk matrix. A custom named-entity recognition model, build from fine tuning a pre-trained language model, is used to identify failure modes, failure causes, failure effects, control processes, and recommendations to aid in failure modes and effects analysis (FMEA). Throughout the presentation, we discuss and apply natural language processing techniques to better leverage the vast amount of information contained in report repositories.

Machine learning↗

NGEE Arctic Field-to-Model

This repository contains workshop documentation and shell scripts/computational infrastructure for building, running, and analyzing output from the DOE Energy Exascale Earth System Model (E3SM) Land Model (ELM).

Fiorella, Rich [@lanl]↗

Dataset Repository for Investigating Suicide Risk Using Social and Environmental Determinants of Health

Suicide is frequently modeled as a function of genetics and environment, where the latter refers to factors other than direct biological consequences, such as air quality, financial level, social connectivity, transportation and food access, and homelessness status. According to the World Health Organization, clean air, a stable climate, adequate water, sanitation and hygiene, safe chemical use, radiation protection, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved natural environment are all prerequisites for good health. Understanding the relationships between these determinants and mental health outcomes requires standardized data that can be included in healthcare programs and health outcome models. There is a wealth of publicly available data on social and environmental factors provided by various US organizations that can benefit the design of health care systems and public health interventions, as well as improve our comprehension of factors that impact health. Such information would not only help improve the understanding of individual and community risk but also identify new risk factors that have not previously been therapeutically targeted, especially in terms of their impact on mental health. However, curating and standardizing such datasets is challenging because they are often recorded at numerous geographical and temporal resolutions and with varying spatial and temporal granularities. To address this challenge, we launched an endeavor in conjunction with the Veterans Health Administration to collect publicly available socioeconomic and environmental determinants of health statistics in the US. In this manuscript, we describe a social and environmental determinants of health (SEDH) datasets repository, data curation documentation, and a pipeline framework for data generation; This effort started in 2020, when we began constructing a scalable pipeline to automate the download, extraction, preparation, analysis, and production of datasets. These datasets have been made available to the VHA and may be shared upon agreement with collaborating organizations.

60 APPLIED LIFE SCIENCES↗