Search NASA⌕ Search

SEARCH · Search NASA

Results for “system modeling ontologies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Automatic determination of fault effects on aircraft functionality

The problem of determining the behavior of physical systems subsequent to the occurrence of malfunctions is discussed. It is established that while it was reasonable to assume that the most important fault behavior modes of primitive components and simple subsystems could be known and predicted, interactions within composite systems reached levels of complexity that precluded the use of traditional rule-based expert system techniques. Reasoning from first principles, i.e., on the basis of causal models of the physical system, was required. The first question that arises is, of course, how the causal information required for such reasoning should be represented. The bond graphs presented here occupy a position intermediate between qualitative and quantitative models, allowing the automatic derivation of Kuipers-like qualitative constraint models as well as state equations. Their most salient feature, however, is that entities corresponding to components and interactions in the physical system are explicitly represented in the bond graph model, thus permitting systematic model updates to reflect malfunctions. Researchers show how this is done, as well as presenting a number of techniques for obtaining qualitative information from the state equations derivable from bond graph models. One insight is the fact that one of the most important advantages of the bond graph ontology is the highly systematic approach to model construction it imposes on the modeler, who is forced to classify the relevant physical entities into a small number of categories, and to look for two highly specific types of interactions among them. The systematic nature of bond graph model construction facilitates the process to the point where the guidelines are sufficiently specific to be followed by modelers who are not domain experts. As a result, models of a given system constructed by different modelers will have extensive similarities. Researchers conclude by pointing out that the ease of updating bond graph models to reflect malfunctions is a manifestation of the systematic nature of bond graph construction, and the regularity of the relationship between bond graph models and physical reality.

Feyock, Stefan↗

FAD-Toolset (Floating Array Design Toolset) [SWR-26-056]

The Floating Array Design (FAD) Toolset is a collection of tools for modeling and designing arrays of floating offshore structures. It was originally designed for floating wind systems but has applicability for many offshore applications. A core part of the FAD Toolset is the floating array model, which serves as a high-level library for efficiently modeling a floating array, such as a floating wind array. It combines site condition information and a description of the floating array design, and contains functions for evaluating the array's behavior considering the site conditions. For example, it combines information about site soil conditions, mooring line loads, and an array's anchor characteristics to estimate the holding capacity of each anchor. The library works in conjunction with the tools RAFT, MoorPy, and FLORIS to model floating platforms, wind turbines, mooring systems, power cables, and array wakes respectively. Layered on top of the floating array model is a set of design tools that can be used for algorithmically adjusting or optimizing parts of the a floating array. Specific tools existing for mooring lines, shared mooring systems, dynamic power cables, static power cable routing, and overall array layout. These capabilities work with the design representation and evaluation functions in the floating array model, and they can be applied by users in various combinations to suit different purposes. In addition to standalone uses of the FAD Toolset, a coupling has been made with Ard, (https://github.com/NLRWindSystems/Ard) a sophisticated and flexible wind farm optimization tool. This coupling allows Ard to use certain mooring system capabilities from FAD to perform layout optimization of floating wind farms with Ard's more advanced layout optimization capabilities. The FAD Toolset works with the IEA Wind Task 49 Ontology (https://github.com/IEAWindTask49/Ontology), which provides a standardized format for describing floating wind farm sites and designs. See example use cases in our examples folder (https://github.com/NLRWindSystems/FAD-Toolset/blob/main/examples/README.md) For working with the library, it is important to understand the floating array model structure, which is described more here: https://github.com/NLRWindSystems/FAD-Toolset/blob/main/fad/README.md.

Sirkis, Leah [National Laboratory of the Rockies (↗

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Theory and ontology for sharing temporal knowledge

Using current technology, the sharing or re-using of knowledge-bases is very difficult, if not impossible. ARPA has correctly recognized the problem and funded a knowledge sharing initiative. One of the outcomes of this project is a formal language called Knowledge Interchange Format (KIF) for representing knowledge that could be translated into other languages. Capturing and representing design knowledge and reasoning with them have become very important for NASA who is a pioneer of innovative design of unique products. For upgrading an existing design for changing technology, needs, or requirements, it is essential to understand the design rationale, design choices, options and other relevant information associated with the design. Capturing such information and presenting them in the appropriate form are part of the ongoing Design Knowledge Capture project of NASA. The behavior of an object and various other aspects related to time are captured by the appropriate temporal knowledge. The captured design knowledge will be represented in such a way that various groups of NASA who are interested in various aspects of the design cycle should be able to access and use the design knowledge effectively. To facilitate knowledge sharing among these groups, one has to develop a very well defined ontology. Ontology is a specification of conceptualization. In the literature several specific domains were studied and some well defined ontologies were developed for such domains. However, very little, or no work has been done in the area of representing temporal knowledge to facilitate sharing. During the ASEE summer program, I have investigated several temporal models and have proposed a theory for time that is flexible to accommodate the time elements, such as, points and intervals, and is capable of handling the qualitative and quantitative temporal constraints. I have also proposed a primitive temporal ontology using which other relevant temporal ontologies can be built. I have investigated various issues of sharing knowledge and have proposed a formal framework for modeling the concept of knowledge sharing. This work may be implemented and tested in the software environment supplied by Knowledge Based System, Inc.

Loganantharaj, Rasiah↗

Scientific Digital Libraries, Interoperability, and Ontologies

Scientific digital libraries serve complex and evolving research communities. Justifications for the development of scientific digital libraries include the desire to preserve science data and the promises of information interconnectedness, correlative science, and system interoperability. Shared ontologies are fundamental to fulfilling these promises. We present a tool framework, some informal principles, and several case studies where shared ontologies are used to guide the implementation of scientific digital libraries. The tool framework, based on an ontology modeling tool, was configured to develop, manage, and keep shared ontologies relevant within changing domains and to promote the interoperability, interconnectedness, and correlation desired by scientists.

Information Model↗

ATOS-1: Designing the infrastructure for an advanced spacecraft operations system

The space industry has identified the need to use artificial intelligence and knowledge based system techniques as integrated, central, symbolic processing components of future mission design, support and operations systems. Various practical and commercial constraints require that off-the-shelf applications, and their knowledge bases, are reused where appropriate and that different mission contractors, potentially using different KBS technologies, can provide application and knowledge sub-modules of an overall integrated system. In order to achieve this integration, which we call knowledge sharing and distributed reasoning, there needs to be agreement on knowledge representations, knowledge interchange-formats, knowledge level communications protocols, and ontology. Research indicates that the latter is most important, providing the applications with a common conceptualization of the domain, in our case spacecraft operations, mission design, and planning. Agreement on ontology permits applications that employ different knowledge representations to interwork through mediators which we refer to as knowledge agents. This creates the illusion of a shared model without the constraints, both technical and commercial, that occur in centralized or uniform architectures. This paper explains how these matters are being addressed within the ATOS program at ESOC, using techniques which draw upon ideas and standards emerging from the DARPA Knowledge Sharing Effort. In particular, we explain how the project is developing an electronic Ontology of Spacecraft Operations and how this can be used as an enabling component within space support systems that employ advanced software engineering. We indicate our hope and expectation that the core ontology developed in ATOS, will permit the full development of standards for such systems throughout the space industry.

Poulter, K. J.↗

BrickQA: Bridging the Semantic Gap in Building Operations with Dynamic Graph Exploration

While standardized ontologies like the Brick schema address data heterogeneity in Building Automation Systems (BAS), accessing this semantic data remains a challenge as domain experts often lack the expertise to formulate complex SPARQL queries. To bridge this gap, we present BrickQA, a Large Language Model (LLM)-based framework that translates natural language into executable SPARQL queries through structured query decomposition, dynamic schema exploration, and inline validation. BrickQA utilizes an iterative reasoning agent to actively navigate graph topology through dynamic exploration actions without requiring exhaustive context injection or model fine-tuning. This approach effectively mitigates hallucinations, particularly in large-scale building knowledge graphs. Empirical evaluation on BuildingQA, a standardized benchmark, demonstrates that BrickQA significantly outperforms ReAct baselines, delivering a 0.291–0.355 absolute F1 improvement while achieving 3 × –12.7 × higher token cost-efficiency. Beyond these metrics, the framework maintains structural fidelity across heterogeneous buildings and remains resilient to ambiguous queries without requiring site-specific fine-tuning. Furthermore, a case study on operational analytics validates the framework’s capability to handle temporal and aggregation constraints, effectively transforming abstract semantic models into actionable facility management insights.1

Ko, Yun-Dam↗

Expert Seeker: A People-Finder Knowledge Management System

The first objective for this report was to perform a comprehensive research of industry models currently being used for similar purposes, in order to provide the Center with ideas of what is being done in area by private companies and government agencies. The second objective was to evaluate the use of taxonomies or ontologies to describe and catalog the areas of expertise at GSFC. The creation of a knowledge taxonomy is necessary for information extraction in order for The Expert Seeker to adequately search and find experts in a particular area of expertise. The requirements to develop a taxonomy are: provide minimal descriptive text; have the appropriate level of abstration; facilitate browsing; ease of use and speed of data entry are critical for success; customized to the organization and its culture; extent of knowledge areas; expandable, so new skills could be develop; could be complemented with free text fields to allow users the option to describe their knowledge in detail.

Becerra-Fernandez, Irma↗

Enabling Interoperability in Earth System Digital Twins (ESDT): Integrating Observations, Models, and AI for Actionable Insights Through NASA'S Intelligent Systems Technology Program

NASA’s Intelligent Systems Technology Program (IST) is driving a paradigm shift in Earth science through the development of Earth System Digital Twins (ESDT). These integrated information systems create a dynamic "digital replica" of the Earth by harmonizing continuous, multi-source observations with high-fidelity models and state-of-the-art artificial intelligence (AI) that enable “What now?”, “What next?”, and “What if?” scenario building. These scenarios are reflected in NASA IST’s series of ESDTs, from the Coastal Zone Digital Twin that integrates complex data on the current state of the Chesapeake Bay to the Terrestrial Environmental Rapid-Replication and Assimilation Hydrometeorological (TerraHydro) AI-based ESDT that forecasts water movement across Earth’s surface, to the Agriculture Land Information System (AgLIS) which can be used to assess optimal planting dates and crop yield estimates. By bridging the gap between vast data archives and actionable insights, these projects enable a system-of-systems approach to understanding complex, interacting Earth processes. This poster will highlight recent innovations and future directions from NASA’s ESDT initiatives: Continuous Data Assimilation & Multi-Source Fusion. A core requirement of the ESDT work is the transition from static models to dynamic "living" replicas. This involves creating frameworks for the continual assimilation of near-real-time data from uncoordinated, heterogeneous sources, including satellite observations and airborne assets, and ground-based Internet of Things (IoT) sensors. These systems link design, operational status, and environmental data, ensuring the digital twin accurately reflects the current state of the physical Earth system. High-Fidelity Hybrid Modeling & Computational Acceleration to enable interactive "what-if" explorations, programs are moving beyond traditional, slow physical solvers by developing fast surrogate machine learning models and Deep Generative Models (DGMs). These hybrid approaches use neural networks to emulate complex physics, such as cloud feedback or ocean dynamics, at a fraction of the original computing cost, often leveraging advanced hardware like Graphics Processing Units (GPUs) to achieve the necessary scale. Federated Ecosystems & Interoperable Frameworks rather than building isolated tools, NASA IST is moving toward federated ESDTs and reusable analytic collaborative frameworks. This theme focuses on interoperability standards and common ontologies that allow specialized digital twins to interact and share data. This system-of-systems architecture supports multi-discipline investigations, such as analyzing how upstream watershed changes impact downstream urban flooding or how wildfire emissions affect regional air quality. By leveraging these advancements, ESDTs empower researchers and decision-makers to conduct real-time analysis and run complex hypothetical scenarios, ultimately improving our understanding of Earth’s evolving systems and informing critical real-world applications.

Earth System↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Information Extraction for System-Software Safety Analysis: Calendar Year 2008 Year-End Report

This annual report describes work to integrate a set of tools to support early model-based analysis of failures and hazards due to system-software interactions. The tools perform and assist analysts in the following tasks: 1) extract model parts from text for architecture and safety/hazard models; 2) combine the parts with library information to develop the models for visualization and analysis; 3) perform graph analysis and simulation to identify and evaluate possible paths from hazard sources to vulnerable entities and functions, in nominal and anomalous system-software configurations and scenarios; and 4) identify resulting candidate scenarios for software integration testing. There has been significant technical progress in model extraction from Orion program text sources, architecture model derivation (components and connections) and documentation of extraction sources. Models have been derived from Internal Interface Requirements Documents (IIRDs) and FMEA documents. Linguistic text processing is used to extract model parts and relationships, and the Aerospace Ontology also aids automated model development from the extracted information. Visualizations of these models assist analysts in requirements overview and in checking consistency and completeness.

Malin, Jane T.↗

Automation for System Safety Analysis

This presentation describes work to integrate a set of tools to support early model-based analysis of failures and hazards due to system-software interactions. The tools perform and assist analysts in the following tasks: 1) extract model parts from text for architecture and safety/hazard models; 2) combine the parts with library information to develop the models for visualization and analysis; 3) perform graph analysis and simulation to identify and evaluate possible paths from hazard sources to vulnerable entities and functions, in nominal and anomalous system-software configurations and scenarios; and 4) identify resulting candidate scenarios for software integration testing. There has been significant technical progress in model extraction from Orion program text sources, architecture model derivation (components and connections) and documentation of extraction sources. Models have been derived from Internal Interface Requirements Documents (IIRDs) and FMEA documents. Linguistic text processing is used to extract model parts and relationships, and the Aerospace Ontology also aids automated model development from the extracted information. Visualizations of these models assist analysts in requirements overview and in checking consistency and completeness.

Malin, Jane T.↗

AI-Powered Knowledge Graphs for Neuromorphic and Energy-Efficient Computing

The surge in scientific literature obscures breakthroughs and hinders the discovery of new research paths. We propose an artificial intelligence (AI) powered framework using large language models (LLMs) and knowledge graphs (KGs) to automate parts of scientific discovery, focusing on energy-efficient AI circuits. Our hybrid approach combines LLMs, structured data, and ontology-based reasoning to construct a comprehensive knowledge graph that integrates insights across computational neuroscience, spiking neuron models, learning rules, architectural motifs, and neuromorphic device technologies. This multi-domain representation enables the generation of hypotheses that connect biological function with implementable, energy-efficient hardware architectures. Using KG embeddings and graph neural networks, the framework generates hypotheses for novel circuits, validates them through optimization on exascale HPC systems, and with tools like SuperNeuro and Fugu, the most promising designs will be prototyped in hardware. This open-source system aims to accelerate discoveries and bridging neuroscience with hardware innovation, drive collaboration, and unlock new opportunities in low-power AI computing.

Gautam, Ashish [ORNL]↗

A semantics-driven framework to enable demand flexibility control applications in real buildings

Decarbonising and digitalising the energy sector requires scalable and interoperable Demand Flexibility (DF) applications. Semantic models are promising technologies for achieving these goals, but existing studies focused on DF applications exhibit limitations. These include dependence on bespoke ontologies, lack of computational methods to generate semantic models, ineffective temporal data management and absence of platforms that use these models to easily develop, configure and deploy controls in real buildings. This paper introduces a semantics-driven framework to enable DF control applications in real buildings. The framework supports the generation of semantic models that adhere to Brick and SAREF while using metadata from Building Information Models (BIM) and Building Automation Systems (BAS). The work also introduces a web platform that leverages these models and an actor and microservices architecture to streamline the development, configuration and deployment of DF controls. The paper demonstrates the framework through a case study, illustrating its ability to integrate diverse data sources, execute DF actuation in a real building, and promote modularity for easy reuse, extension, and customisation of applications. The paper also discusses the alignment between Brick and SAREF, the value of leveraging BIM data sources, and the framework's benefits over existing approaches, demonstrating a 75% reduction in effort for developing, configuring, and deploying building controls.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure

Modern experimental platforms such as particle accelerators, fusion devices, telescopes, and industrial process control systems expose tens to hundreds of thousands of control and diagnostic channels, accumulated over decades of hardware evolution. Operators and AI systems alike depend on informal expert knowledge, inconsistent naming conventions, and scattered documentation to locate the signals required for monitoring, troubleshooting, and automated control, creating a persistent bottleneck for reliability, scalability, and emerging language-model-driven interfaces. We formalize semantic channel finding, the task of mapping natural-language intent to concrete control-system signals, as a general problem in complex experimental infrastructure, and introduce a four-paradigm conceptual framework to guide architecture selection based on facility-specific data regimes. The paradigms span (i) direct in-context lookup over small, curated channel dictionaries, (ii) constrained hierarchical navigation through structured trees, (iii) interactive agent exploration using iterative reasoning and tool-based database queries, and (iv) ontology-grounded semantic search that decouples channel meaning from facility-specific naming conventions. We demonstrate the practical feasibility of each paradigm through proof-of-concept implementations at four operational facilities spanning two orders of magnitude in scale: from compact free-electron lasers to large synchrotron light sources, operating under diverse control-system architectures ranging from clean hierarchical naming schemes to legacy environments with decades of heterogeneous conventions. Where evaluated against expert-curated operational queries, these instantiations achieve 90%–97% accuracy, validating the framework’s applicability across real-world deployment scenarios. To accelerate adoption across the broader scientific and industrial control-system community, we release open-source, plug-and-play implementations of all three interactive paradigms-direct lookup, hierarchical navigation, and middle-layer exploration-within the Osprey framework, together with tools for channel database generation, interactive testing, and minimal-configuration deployment. This work establishes semantic channel finding as a foundational capability for human-centric and agentic AI interfaces at large-scale facilities, providing both a systematic framework for architecture design and practical resources to enable adoption without building custom infrastructure from scratch.

channel finding↗

Enabling Assurance in the MBSE Environment

A number of specific benefits that fit within the hallmarks of effective development are realized with implementation of model-based approaches to systems and assurance. Model Based Systems Engineering (MBSE) enabled by standardized modeling languages (e.g., SysML®) is at the core. These benefits in the context of spaceflight system challenges can include [1]: • Improved management of complex development • Reduced risk in the development process • Improved cost management • Improved design decisions With appropriate modeling techniques the assurance community also can improve early oversight and insight into project development. NASA has shown the basic constructs of SysML in an MBSE environment offer several key advantages, within a Model Based Mission Assurance (MBMA) initiative [2, 3]. These include the following: • Model viewpoints that promote rapid and systematic assessment of requirements coverage, hazard tagging and risk management • Embedded safety assessments for launch vehicles • Deployment of model assisted development of reliability products - Failure Modes and Effects Analyses (FMEAs) and Fault Trees • Test Planning • Validation and Verification of complex functions • Support of Assurance Case development for complex systems In addition, while there are benefits to be harvested, there is a realization that these do not come without effort and cost. Enabling model-based approaches requires structure, not only in an organizational context, but in a modeling context as well. There can be a steep learning curve and costs associated to train skilled modelers. But, on the other hand, not all of the assurance community need to be modelers. Models themselves must conform to ontologies that enable assurance. This places constraints upon the models and modelers. Optimums have yet to be developed where resources and constraints on modeling must be traded off in the organization and modeling efforts for projects. A number of barriers need to be overcome, as well, which pose challenges to the developers of the software that supports MBSE/MBMA. Information and data must be made to flow seamlessly through the life cycle. Because there is a wide variety of tools used in the community, to avoid the problems of the past of silos, delays, and diverging interests, information should flow among these tools to support the “single source of truth” paradigm of MBSE. This will greatly facilitate MBMA and advancement of assurance functions.

Evans, John W.↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗