Search NASA⌕ Search

SEARCH · Search NASA

Results for “Ontology engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ontology Development Kit: a toolkit for building, maintaining and standardizing biomedical ontologies

Similar to managing software packages, managing the ontology life cycle involves multiple complex workflows such as preparing releases, continuous quality control checking and dependency management. To manage these processes, a diverse set of tools is required, from command-line utilities to powerful ontology-engineering environmentsr. Particularly in the biomedical domain, which has developed a set of highly diverse yet inter-dependent ontologies, standardizing release practices and metadata and establishing shared quality standards are crucial to enable interoperability. The Ontology Development Kit (ODK) provides a set of standardized, customizable and automatically executable workflows, and packages all required tooling in a single Docker image. In this paper, we provide an overview of how the ODK works, show how it is used in practice and describe how we envision it driving standardization efforts in our community.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Simplifying and Visualizing the Ontology of Systems Engineering Models

The credibility of an engineering model is of critical importance in large-scale projects. How concerned should an engineer be when reusing someone else's model when they may not know the author or be familiar with the tools that were used to create it? In this report, the authors advance engineers' capabilities for assessing models through examination of the underlying semantic structure of a model--the ontology. This ontology defines the objects in a model, types of objects, and relationships between them. In this study, two advances in ontology simplification and visualization are discussed and are demonstrated on two systems engineering models. These advances are critical steps toward enabling engineering models to interoperate, as well as assessing models for credibility. For example, results of this research show an 80% reduction in file size and representation size, dramatically improving the throughput of graph algorithms applied to the analysis of these models. Finally, four future problems are outlined in ontology research toward establishing credible models--ontology discovery, ontology matching, ontology alignment, and model assessment.

42 ENGINEERING↗

Dynamic Retrieval Augmented Generation of Ontologies using Artificial Intelligence (DRAGON-AI)

Ontologies are fundamental components of informatics infrastructure in domains such as biomedical, environmental, and food sciences, representing consensus knowledge in an accurate and computable form. However, their construction and maintenance demand substantial resources and necessitate substantial collaboration between domain experts, curators, and ontology experts. We present Dynamic Retrieval Augmented Generation of Ontologies using AI (DRAGON-AI), an ontology generation method employing Large Language Models (LLMs) and Retrieval Augmented Generation (RAG). DRAGON-AI can generate textual and logical ontology components, drawing from existing knowledge in multiple ontologies and unstructured text sources.We assessed performance of DRAGON-AI on de novo term construction across ten diverse ontologies, making use of extensive manual evaluation of results. Our method has high precision for relationship generation, but has slightly lower precision than from logic-based reasoning. Our method is also able to generate definitions deemed acceptable by expert evaluators, but these scored worse than human-authored definitions. Notably, evaluators with the highest level of confidence in a domain were better able to discern flaws in AI-generated definitions. We also demonstrated the ability of DRAGON-AI to incorporate natural language instructions in the form of GitHub issues.These findings suggest DRAGON-AI's potential to substantially aid the manual ontology construction process. However, our results also underscore the importance of having expert curators and ontology editors drive the ontology generation process.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Model Assessment Wizard (MAW): A Visualization System for Ontology Constraint Violations

Validation and verification of engineering models is important to understand potential weaknesses and issues in the model. This is accomplished through the application of constraint logic to the model. These models and the constraints put upon them can be represented through a graph structure. Here we give a visualization system to aid users understanding, locating, and fixing constraint violations in their systems. We give users several ways to narrow down on the specific errors and parts of the graph they’re interested in. Users have the opportunity to choose the types of errors that will be shown in the graph. Clustering is applied to the graph to help users narrow down their searches. Several other graph interactions are given to support discovery of constraint violations.

97 MATHEMATICS AND COMPUTING↗

Retaining Systems Engineering Model Meaning Through Transformation: Demo 2

Digital engineering strategies typically assume that digital engineering models interoperate seamlessly across the multiple different engineering modeling software applications involved, such as model- based systems engineering (MBSE), mechanical computer-aided design (MCAD), electrical computer-aided design (ECAD), and other engineering modeling applications. The presumption is that the data schema in these modeling software applications are structured in the familiar flat- tabular schema like any other software application. Engineering domain-specific applications (e.g., systems, mechanical, electrical, simulation) are typically designed to solve domain-specific problems, necessarily excluding explicit representations of non-domain information to help the engineer focus on the domain problems (system definition, design, simulation). Such exclusions become problematic in inter-domain information exchange. The obvious assumptions of one domain might not be so obvious to experts in another domain. Ambiguity in domain-specific language can erode the ability to enable different domain modeling applications to interoperate, unless the underlying language is understood and used as the basis for translation from one application to another. The engineering modeling software application industry has struggled for decades to enable these applications to interoperate. Industry standards have been developed, but they have not unified the industry. Why is this? The authors assert that the industry has relied on traditional database integration methods. The basic issue prohibiting successful application integration then is that traditional database-driven integration does not consider the distinct languages of each domain. An engineering models meaning is expressed through the underlying language of that engineering domain. In essence, traditional integration methods do not retain the semantic context (meaning) of the model. The basis of this research stems from the widely held assumption that systems engineering models are (or can be) structured according to the underlying semantic ontology of the model. This assumption can be imagined from two thoughts. 1) Digital systems engineering models are often represented using graph theory (the graph of a complex systems model can contain millions of nodes and edges). When examining the nodes one at a time and following the outbound edges of each node one by one, one can end up with rudimentary statements about the model (i.e., node A relates to node B), as in a semantic graph. 2) Likewise, from the study of natural languages, a sentence can be structured into unambiguous triples of subject-predicate-object within formal and highly expressive semantic ontologies. The rudimentary statements about a systems model discerned with graph theory closely mimic the triples used in the ontologies that try to structure natural languages. In other words, a systems models semantic graph can be (or is) structured into an ontology. Additionally, it is well established in industry that through natural language processing (NLP), which provides the means to create language structures, that computers can interpret ontological graphs. Therefore, the authors hypothesized that if the integrity of the underlying semantic structure of a systems model is retained, the contextual meaning of the model is retained. By structuring system models into the triples of the underlying ontology during the transformation from one MBSE application to another, the authors have provided a proof of the concept that the meaning of a system model can be retained during transformation. The authors assert that this is the missing ingredient in effective systems model-to-model interoperability. ACKNOWLEDGEMENTS The authors would like to thank the FY19 Model Interoperability team members who provided a solid foundation for the FY20 team to leverage: John McCloud, for the work he did to guide us toward the right use of technology that will appropriately discover and manipulate ontologies. Carlos Tafoya, for the work he did to develop an application programming interface (API)/Adapter that would export ontology-based data from GENESYS. Peter Chandler, for the work he did to architect our overall integration solution, with an eye toward the future that would influence a large-scale federated production-level systems engineering digital model ecosystem.

42 ENGINEERING↗

TriGORank: A Gene Ontology Enriched Learning-to-Rank Framework for Trigenic Fitness Prediction

Machine learning (ML) has been gaining interest in the metabolic engineering community as a means to automate prediction tasks. In this work, we introduce and study the task of using ML to recommend high-fitness triplet mutants as candidates for wet-lab experiments. We first utilize individual fitness and digenic fitness scores as features and train machine learning models that produce a ranked list, from high to low fitness scores, for triplet gene mutants of S. cerevisiae. Then, we incorporate prior metabolic knowledge from an existing gene ontology, by designing a novel graph representation and deducing features that can capture gene similarity and gene interactions. Lastly, experimental results show that our proposed gene ontology enriched model, termed TriGORank, improves both performance and explainability.

Labhishetty, Sahiti↗

Improving the User Interface of the DeepLynx Data Warehouse

DeepLynx is an open-source ontology-based data warehouse created by INL to support the creation and life cycle of digital engineering projects, with a particular emphasis on digital twins [1]. Digital twins are systems that represent physical assets and process in a real-time digital environment [1]. Most well-known commercial data warehouses use Graphical User Interfaces (GUIs) for users to interact with their systems [3]. Limited publications have addressed the design of these interfaces and understanding of their target users. The current users and development team acknowledge the need to improve the current UI, not just for aesthetics but to improve functionality and workflow of DeepLynx. Traditional data warehouse users are developers, data scientists and business analysts [2]. DeepLynx users have a vast range of experience using data warehouses, and diverse roles, including engineers, scientists and management positions. Because there is a broader audience of target users for DeepLynx than a typical data warehouse, it is essential that DeepLynx has a useable and intuitive user interface. To achieve this the team performed human-computer interaction methods, including a Heuristic Evaluation of current UI using Neilsen’s Usability Heuristic, create personas based on current users by designing a user survey, data analysis and develop of personas. Followed by a redesign of the UI following using Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design in industry standard software Figma. Lastly a Heuristic Evaluation of new UI design, using Neilsen’s Usability Heuristic and User testing of redesign UI and have a group of users complete a Thinking Aloud Test of the new UI. Preliminary results of the Heuristic Evaluation of current UI arise issue with Consistency and Standards, Visibility of System Status, Match System and Real World and Recognition Rather than Recall. These issues were addressed in the proposed redesign by applying Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design. Next steps include formalized list of lessons learned and design implications for future publications.

97 MATHEMATICS AND COMPUTING↗

Model Assessment Wizard (MAW)

SAND2026-18710O The Model Assessment Wizard (MAW) is a tool for evaluating ontologies and provides users with a comprehensive workbench for analysis. MAW features sub-modules for visualization, alignment, Shapes Constraint Language (SHACL) and Web Ontology Language (OWL) constraints, and simplification. Users can upload data, identify missing information, visualize ontologies, and update constraints. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Murdock, Jaimie [Sandia National Lab. (SNL-CA), Li↗

Visual Brick model authoring tool for building metadata standardization

In this study, the Brick ontology is a unified semantic metadata standard for building assets and their relationships, serving as a key enabler for effective interoperability and automation of building systems and analytics. However, creating a Brick model, in other words, standard semantic metadata based on the Brick ontology for a building dataset, can be a complex task. This paper presents two case studies of the creation of Brick models for real-world residential and commercial building datasets, highlighting the challenges during the Brick model creation process. Additionally, the paper introduces VizBrick, an interactive authoring tool for creating semantic building metadata. VizBrick facilitates the creation of Brick models by providing an intuitive visual interface and interactive capabilities, such as keyword search, automatic mapping suggestions, and recommendations. The use of VizBrick is shown to significantly reduce the time and effort required during the Brick model creation process.

42 ENGINEERING↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

Metadata Schemas and Ontologies for Building Energy Applications: A Critical Review and Use Case Analysis

With the increasing digitalization of processes throughout the lifecycle of buildings, data exchanged between stakeholders and between building systems has grown significantly. However, a lack of semantic interoperability between data in different systems is still prevalent, hindering the development of applications that can be reused across buildings and limiting the scalability of innovative solutions. Semantics refers to the description of the meaning of the data in a way that can be consistently understood by applications. Recently, several competing initiatives have been developing metadata schemas and ontologies to express this semantic information for different applications in the building domain. This paper systematically reviews these schemas and conducts an analysis of five of them to evaluate their applicability to three high-value use cases for building operations: energy audits, automated fault detection and diagnostics and optimal control. The survey finds 40 schemas published in the last 10 years but but their actual use in industry is difficult to estimate. Among the five selected ontologies, several gaps are highlighted in relation to the three use cases. Recommendations for the future include better harmonization of these initiatives, more centralized repositories and search engines for these schemas as well as better industry engagement to facilitate their adoption.

Smart Building, Sematic, Metadata, Ontology, Data ↗

FAIR to WISE (F2W) v1.0.0

FAIR to WISE (F2W) is an iterative, large-language model (LLM) driven pipeline that turns unstructured research PDFs into structured, queryable knowledge graphs (KGs). Core features include schema-driven extraction to a LinkML model; full provenance capture; ontology-grounded enrichment (e.g., chemical validation and ChEBI lookup); graph construction to JSON-LD with stable IDs; and KG-RAG question answering with evidence-aware retrieval. The system is engineered for reproducibility and accessibility (open-source Ollama models, temperature=0, NVTX/Nsight profiling) with robust QA (relation verification, deduplication, and deterministic outputs). Primary uses are literature-to-KG automation, knowledge-grounded Q&A, and experimental steering support. We demonstrate the approach in organic photovoltaics, where the pipeline ingests papers, builds a domain KG, and evaluates answers against expert competency questions to guide experimental planning and interpretation. Compared with off-the-shelf LLMs and ad-hoc NLP tools, F2W addresses ontology gaps and reduces hallucination risk by grounding responses in extracted evidence and enforcing schema constraints; it also offers deterministic, provenance-linked outputs and open, cost-aware deployment. Evidence-aware ranking further improves answer quality over pure vector search.

Abramov, David [Lawrence Berkeley National Laborat↗

Development of a Framework for Data Integration, Assimilation, and Learning for Geological Carbon Sequestration (DIAL-GCS) (Final Report)

This project aimed to develop and demonstrate a Data Integration, Assimilation, and Learning framework for geologic carbon sequestration projects (DIAL-GCS). DIAL-GCS is an intelligence monitoring system (IMS) for automating GCS closed-loop management by leveraging recent developments in machine learning technologies, complex event processing (CEP), and reduced-order modeling. The safe and efficient operation of GCS repositories requires integrated monitoring to track the injected CO¬2 as it moves within a storage reservoir. GCS projects are data intensive, as a result of proliferation of digital instrumentation and smart-sensing technologies. GCS projects are also resource intensive, often requiring multidisciplinary teams performing different monitoring, verification, accounting (MVA) tasks throughout the lifecycle of a project to ensure secure containment of injected CO2. The success of GCS thus depends in a large part on our ability to access, assimilate, and analyze heterogeneous data and information sources in a timely manner. This project included a number of meaningful and necessary tasks to transform the human domain knowledge into machine-interpretable rules for automating knowledge extraction and discovery in GCS. The specific technical objectives of the proposed DIAL-GCS project were to develop an ontology-driven GCS data management module for storing, querying, and exchanging GCS data (both historic and live sensor data) from multiple sources and in heterogeneous formats. Incorporate a CEP engine for detecting abnormal situations by seamlessly combining expert knowledge, rule-based reasoning, and machine learning. Enable uncertainty quantification and predictive analytics using a combination of coupled-process modeling, AI/ML methods, and reduced-order modeling, and integrate and demonstrate the system’s capabilities with both real and simulated data. As far as we know, this is one of the first projects aimed to develop intelligent monitoring systems (IMS) targeting the GCS. Under this project, the team had developed a large number of web applications and scientific algorithms that contribute the main theme of intelligent monitoring. The team has published more than a dozen peer reviewed papers and disseminated the research results at multiple technical meetings.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A Power Application Developer’s Guide to the Common Information Model: An Introduction for Power Systems Engineers and Application Developers – CIM17v40

A key issue in creating the next generation of energy management system (EMS) and advanced distribution management system (ADMS) platforms will be the ability to represent and exchange power system network model data in a consistent manner. To this end, the Common Information Model (CIM) stands out as the only standardized vocabulary (or ontology) for defining power system network models and asset data in a comprehensive, consistent manner across the generation-transmission-distribution boundary. The CIM is freely available to use and extend. The CIM is maintained by the UCAiug (informally known as the CIM User’s Group) under an Apache 2.0 license. The CIM Users Group collaborates with the IEC and other standards communities for the development of technical and informative specifications. Although portions of the information model are referred to by the corresponding IEC standards naming, it is not necessary to purchase any of the IEC standards to use the CIM information model. This document provides a roadmap for power system engineers and application developers not familiar with semantic modeling to start using the CIM for modeling, simulation, optimization, and development of advanced power applications. The key classes needed for defining power system topology and equipment are explained systematically. Key focus areas include modeling of lines, transformers, generators, switching equipment, loads, and distributed energy resources (DERs).

97 MATHEMATICS AND COMPUTING↗

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

Integrated system failure analysis software toolchain (IS-FAST)

Systems and methods are provided for generating faults and analyzing fault propagation and its effects. Starting from the ontologies of components, functions, flows, and faults, systems and methods are provided that describe, generate and track faults in a computer system across multiple domains throughout design and/or development. In order to construct the system and fault models, a series of concepts is introduced in the form of ontologies and their dependencies. An investigation is performed into the faults, including their type, cause, life-cycle aspects, and effect. Principles and rules are created to generate various faults based on system configurations. After the modeling process, a simulation engine is described to execute actions and simulate the process of fault generation and propagation. As a result, fault paths that impact components and functions can be obtained.

Diao, Xiaoxu↗