Search NASA⌕ Search

SEARCH · Search NASA

Results for “Interoperability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

Investigating Neighbor Discovery for High-Rate Delay Tolerant Networking

A neighbor discovery protocol that enables the dynamic discovery of nodes for NASA Glenn Research Center’s High Rate Delay Tolerant Networking (HDTN) project is needed to satisfy mission scalability and interoperability objectives. The purpose of this study is to primarily examine a DTN IP neighbor discovery (IPND) implementation to understand the inner workings of the protocol before approaching a neighbor discovery implementation for the HDTN project. Using a rust-based DTN IPND implementation, this study analyzed the neighbor discovery process. DTN IPND was found to help meet both mission scalability and interoperability objectives by enabling dynamic routing capabilities. This study definitively answers the question regarding mission scalability and interoperability concerns. Further studies are needed to establish the most optimal neighbor discovery implementation for HDTN.

Prash Choksi↗

Flexible User Radio for Lunar Missions

NASA’s Artemis program and other lunar exploration and development programs are planning over 40 lunar missions before 2030. Lunar missions, both crewed and uncrewed, include orbiters, landers, rovers, and surface stations. All these missions require communications with Earth, either via Direct to Earth (DTE) links or through relays in lunar orbit. Multiple DTE options are available among existing and planned ground stations: Deep Space Network (DSN), European Space Agency, and others. Relay options include the planned Lunar Gateway, LunaNet compliant relays, and some lunar landers propose to launch dedicated orbiters. The dilemma for lunar system designers is to identify a communication link which meets mission requirements but does not have issues of limited access (e.g. DSN is in high demand supporting deep space missions with high priority and some with inflexible schedules), system impacts (high power Radio Frequency (RF) for DTE links), cost (dedicated relay), or operational date. To avoid this difficult decision, a Flexible Radio for Lunar Missions is proposed, which will enable system designs to proceed prior to any final decision on the communication network to be used, by enabling compatibility with any of multiple DTE or orbital relay communication systems. The Flexible Radio will support the necessary frequency, bandwidth, modulation, and power requirements to interoperate with the majority of known or planned DTE or relay systems, and can be designed into a lunar mission without prior knowledge of which link will ultimately be used. Furthermore, the link being used can be changed as needed during the mission, in near real time. The Flexible Radio design will leverage work already completed at NASA in the areas of Wideband RF, Software Defined Radio, Adaptive Coding and Modulation, and Phased Array antennas. The Flexible Radio requires sufficient bandwidth to cover the allocated frequencies for both the operation of links in cislunar space and for space-to-Earth links; the flexibility to support multiple modulations, data rates, and coding schemes; the ability to identify available relays, detect and recognize the signals of those relays, and adapt its own frequency, modulation, symbol rate, and code rate to operate with the detected relay (or DTE station); and finally, it requires appropriate software to support network configuration and interoperability with the detected network. The Flexible Radio can be designed in a sufficiently small, lightweight, and low-power package to be used in a wide variety of lunar systems. The initial implementation, as proposed, will focus on the Ka-band, supporting up to 2 GHz bandwidth around the 27 GHz frequency for return links, and 23 GHz for forward links. Other frequency bands are under consideration for future configurations. The software defined modem will support OQPSK, BPSK, and NASA-defined modulations which also support two-way ranging with data. For near-real time adaptation, the Flexible Radio will scan the sky for available relays, using a phased array antenna with Adaptive/Cognitive Communications. It will then configure for network interoperability supporting DTN and other protocol options. The Flexible Radio will support scheduled connections and on-demand use when available.

Lunar Space Communications↗

NASA’s Safety, Reliability, and Mission Assurance Digital Future

The evolution from “document-centric” to “data-centric” and “model-centric” information leveraging structured data and model-based approaches is at the heart of digital engineering transformational efforts underway across industry and government. It is these approaches that pave the way for data lakes, Authoritative Sources of Truth (ASOTs), and systems- of-systems interoperability and the corresponding transformational benefits thereof. Such benefits include increased data availability, data access equity, data traceability, real-time analytics, batch analytics, and (most importantly) acceleration of the time-to-value and time-to-insights associated with engineering products and analyses. The longer-term benefits of reusability, customization and traceability are even more promising. For Safety and Mission Assurance (SMA), and Mission Success (SMS) activities; realization of such benefits is essential to provide engineers and analysts alike vital information when needed to support critical decision making throughout the entire life cycle. The SMA community often operate in parallel with engineering activities, for which information exchange with relevant context is paramount. Far too often, such information lags key decision points and/or is absent of the robust, integrated, knowledge needed, given inherent barriers associated with traditional document-centric means to data sharing, analysis, and reporting. This paper provides an overview of how NASA’s Office of Safety and Mission Assurance (OSMA) is evolving its policies, standards, guidance, and training to transform to eliminate such barriers, thus realizing the benefits emerging in this new digital era. A roadmap for achieving this digital future is presented along with key building blocks involving use and implementation of concepts such as: Objectives-Hierarchies, Objective-Driven Requirements, Accepted Standards, Safety and Assurance Cases, data digitization (i.e., ontologies, structured data, and model-centric data), FAIR (Findable, Accessible, Interoperable, & Reusable) and/or FAIRUST (Findable, Accessible, Interoperable, Reusable, Understandable, Secure, and Trusted) principles [1]. This paper also describes how OSMA, leveraging the Agency’s overall commitment to Digital Transformation (DT), is using the power of Policy, “Digital” Domain representation, Product Evolution, and Community Outreach and Engagement as part of a strategic vision and roadmap to evolve and transform its SMA organizations to become better able to serve its stakeholders and customers. Future publications will elaborate on these building blocks and deeper concepts.

Authoritative Source of Truth (ASOT),↗

Flexible User Radio for Lunar Missions

NASA’s Artemis program and other lunar exploration and development programs are planning over 40 lunar missions before 2030. Lunar missions, both crewed and uncrewed, include orbiters, landers, rovers, and surface stations. All these missions require communications with Earth, either via Direct to Earth (DTE) links or through relays in lunar orbit. Multiple DTE options are available among existing and planned ground stations: Deep Space Network (DSN), European Space Agency, and others. Relay options include the planned Lunar Gateway, LunaNet compliant relays, and some lunar landers propose to launch dedicated orbiters. The dilemma for lunar system designers is to identify a communication link which meets mission requirements but does not have issues of limited access (e.g. DSN is in high demand supporting deep space missions with high priority and some with inflexible schedules), system impacts (high power Radio Frequency (RF) for DTE links), cost (dedicated relay), or operational date. To avoid this difficult decision, a Flexible Radio for Lunar Missions is proposed, which will enable system designs to proceed prior to any final decision on the communication network to be used, by enabling compatibility with any of multiple DTE or orbital relay communication systems. The Flexible Radio will support the necessary frequency, bandwidth, modulation, and power requirements to interoperate with the majority of known or planned DTE or relay systems, and can be designed into a lunar mission without prior knowledge of which link will ultimately be used. Furthermore, the link being used can be changed as needed during the mission, in near real time. The Flexible Radio design will leverage work already completed at NASA in the areas of Wideband RF, Software Defined Radio, Adaptive Coding and Modulation, and Phased Array antennas. The Flexible Radio requires sufficient bandwidth to cover the allocated frequencies for both the operation of links in cislunar space and for space-to-Earth links; the flexibility to support multiple modulations, data rates, and coding schemes; the ability to identify available relays, detect and recognize the signals of those relays, and adapt its own frequency, modulation, symbol rate, and code rate to operate with the detected relay (or DTE station); and finally, it requires appropriate software to support network configuration and interoperability with the detected network. The Flexible Radio can be designed in a sufficiently small, lightweight, and low-power package to be used in a wide variety of lunar systems. The initial implementation, as proposed, will focus on the Ka-band, supporting up to 2 GHz bandwidth around the 27 GHz frequency for return links, and 23 GHz for forward links. Other frequency bands are under consideration for future configurations. The software defined modem will support OQPSK, BPSK, and NASA-defined modulations which also support two-way ranging with data. For near-real time adaptation, the Flexible Radio will scan the sky for available relays, using a phased array antenna with Adaptive/Cognitive Communications. It will then configure for network interoperability supporting DTN and other protocol options. The Flexible Radio will support scheduled connections and on-demand use when available.

Lunar Space Communications↗

Objective Structured Clinical Evaluation (OSCE) of an Artificial Intelligence (AI) Clinical Decision Support System (CDSS) Tool

BACKGROUND Objective Structured Clinical Evaluations (OSCEs) have long been established as a robust methodology for summative assessment of clinical skills and decision-making during medical education. The recent integration of Artificial Intelligence (AI) into clinical decision-making processes has prompted the need for novel evaluation frameworks to assess the efficacy and reliability of AI clinical decision support system (CDSS) tools. This abstract outlines the process of quantitatively evaluating a novel CDSS (“Doc in a Box” Google 2024) trained on curated medical spaceflight data in the psychomotor domain as it interfaces with a human volunteer acting as the crew medical officer (CMO). PURPOSE The AI CDSS under review was developed as part of the Lunar Command and Control Interoperability (LuCCI) project, which is intended to address a gap in how Lunar Surface Systems (LSS) would interoperate across multiple programs, commercial partners, and international partners. The project objective is to define, prototype, integrate, and evaluate an interoperable lunar command, control, data, and software reference architecture to enable autonomy and informatics capability through common standards across LSS. A multi-modal AI-based CDSS compatible with Federated LSS will assist clinicians in diagnosing and managing complex medical conditions by providing evidence-based recommendations through predictive analytics. Given the critical role of decision-support as NASA continues to evolve its Earth-independent medical operations (EIMO), it is imperative to ensure that such AI tools perform reliably and align with clinical standards during progressive lunar and Martian exploration class missions. METHODS The OSCE framework, traditionally used for evaluating human clinicians, was adapted to assess the AI tool's decision-making capabilities in simulated clinical scenarios. In this adapted OSCE, the AI CDSS was tested across a series of structured clinical scenarios designed to mimic real-life spaceflight patient cases. These scenarios included a range of conditions and complexities, allowing for comprehensive assessment of the tool's performance. Key evaluation metrics included accuracy of diagnosis, timeliness of decision-making, and appropriate recommendations for therapies. The OSCE was scored by human physician evaluators who assessed the AI's recommendations in comparison with expert clinicians' medical decision making to ensure alignment with best practices and the standard of care. RESULTS Preliminary results indicate that the AI CDSS demonstrated high accuracy in diagnostic recommendations and decision support across various scenarios. However, certain limitations were noted, such as occasional discrepancies in handling complex or nuanced cases that required a more contextual understanding. Additionally, the tool scored higher on the diagnostic portion of the rubric, with lower scores in the therapeutic recommendations. These findings highlight the importance of continuous refinement and validation of AI tools through rigorous evaluation frameworks like the OSCE. The adaptation of OSCEs for AI tools presents several advantages, including a structured and reproducible approach to evaluation, the ability to test AI systems in diverse clinical scenarios, and the opportunity to benchmark AI performance against established clinical standards to permit charting of future progress as aerospace medicine evolves as a discipline. Remaining challenges include ensuring that these evaluations capture the full spectrum of clinical decision-making scenarios that will be confronted by CMOs during missions and adequately reflecting real-world variability of the austere spaceflight environment. CONCLUSION Employing OSCEs to evaluate AI clinical decision support tools offers a promising approach to validating their clinical utility and efficacy. This methodology not only provides insights into the tool's performance but also fosters ongoing improvement and alignment with standard of care practices. Future research should focus on refining these evaluation processes and addressing limitations to enhance the integration of AI tools in clinical spaceflight settings. REFERENCES Scott S, Hearns V, Barker MA. Testing Clinical Skills: A Look at the OSCE and USMLE Clinical Skills Exams. S D Med. 2019 Oct;72(10):451-453. Majumder MAA, Kumar A, Krishnamurthy K, Ojeh N, Adams OP, Sa B. An evaluative study of objective structured clinical examination (OSCE): students and examiners perspectives. Adv Med Educ Pract. 2019 Jun 5;10:387-397. Karam VY, Park YS, Tekian A, Youssef N. Evaluating the validity evidence of an OSCE: results from a new medical school. BMC Med Educ. 2018 Dec 20;18(1):313.

Ariana M Nelson↗

Bridging Control and Deployment: A Cross-Layer Analysis of Scalable Building Cluster Control

Building cluster control has emerged as a promising approach for enabling flexible and coordinated operation of distributed building systems, yet its transition from pilot demonstrations to routine grid-interactive operation remains limited. This paper argues that this gap cannot be explained by control algorithms alone. Instead, it arises from interacting barriers in communication infrastructure, data and semantic interoperability, uncertainty management, stakeholder participation, market design, and policy support. Accordingly, the paper reviews both technical and non-technical barriers to building cluster control. Technical challenges include heterogeneous devices and protocols, communication latency and reliability, distributed decision-making, and uncertainty propagation across aggregated loads. Non-technical barriers include user participation, stakeholder coordination, incentive allocation, and data governance. Existing solution approaches are synthesized, including semantic interoperability frameworks, edge and hierarchical communication architectures, distributed and transactive control strategies, uncertainty-aware optimization, policy mechanisms, and market reforms. Based on this analysis, two research directions are identified: testing infrastructures that can evaluate control performance under realistic multi-building conditions, and abstraction methods that allow building clusters to interact with other energy sectors through standardized flexibility representations. Overall, the paper provides a structured review of how building cluster control can move from isolated demonstrations toward reproducible, market-compatible, and grid-relevant implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship. 1. Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. 2. National Academies of Sciences, E. and Medicine, Open Science by Design: Realizing a Vision for 21st Century Research. 2018, Washington, DC: The National Academies Press. 232. 3. Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5.

Life Sciences data↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.

Life Sciences data↗

A Discussion of Time Management Concepts and Time Constraint Equations for Multi-Rate Federation Executions

The High Level Architecture (HLA) is a simulation interoperability standard developed by the Simulation Interoperability Standards Organization (SISO) and published as the international standard IEEE 1516-2010 by the Institute for Electrical and Electronics Engineers (IEEE). HLA is a widely used standard for the development and execution of collaborative distributed simulations. HLA provides a number of Management Services to simulation developers: Federation, Declaration, Object, Ownership, Data Distribution, and Time. Of those services, Time Management Services is probably one of the least understood and least used. However, Time Management Services are critical to technical simulations like those created for space systems using the Space Reference Federation Object Model (SpaceFOM). Time Management can be used to insure data coherence and execution repeatability in distributed simulations. When combined with real time execution policies, Time Management is being used to support real time execution of mixed software and hardware in the loop integration, verification, and validation simulations for active space systems development. This paper starts by providing an overview of the HLA Time Management Services. This provides the background to discuss the challenges associated with Time Management and its use, starting with simple common rate frame scheduled simulations, then simple multi-rate simulations, and ending with complex mixed rate simulations. The authors then formulate the significant time constraint relationships between identified frame scheduling parameters. The intent of the paper is to provide a concise discussion of how to use Time Management in both simple cases and in more complex mixed frame rate federation executions.

Simulation Interoperability↗

A Discussion of Time Management Concepts and Time Constraint Equations for Multi-Rate Federation Executions

The High Level Architecture (HLA) is a simulation interoperability standard developed by the Simulation Interoperability Standards Organization (SISO) and published as the international standard IEEE 1516-2010 by the Institute for Electrical and Electronics Engineers (IEEE). HLA is a widely used standard for the development and execution of collaborative distributed simulations. HLA provides a number of Management Services to simulation developers: Federation, Declaration, Object, Ownership, Data Distribution, and Time. Of those services, Time Management Services is probably one of the least understood and least used. However, Time Management Services are critical to technical simulations like those created for space systems using the Space Reference Federation Object Model (SpaceFOM). Time Management can be used to insure data coherence and execution repeatability in distributed simulations. When combined with real time execution policies, Time Management is being used to support real time execution of mixed software and hardware in the loop integration, verification, and validation simulations for active space systems development. This paper starts by providing an overview of the HLA Time Management Services. This provides the background to discuss the challenges associated with Time Management and its use, starting with simple common rate frame scheduled simulations, then simple multi-rate simulations, and ending with complex mixed rate simulations. The authors then formulate the significant time constraint relationships between identified frame scheduling parameters. The intent of the paper is to provide a concise discussion of how to use Time Management in both simple cases and in more complex mixed frame rate federation executions.

Simulation Interoperability↗

Fragme∩t: An Open‐Source Framework for Multiscale Quantum Chemistry Based on Fragmentation

Fragment-based quantum chemistry offers a means to circumvent the nonlinear computational scaling of conventional electronic structure calculations, by partitioning a large calculation into smaller subsystems then considering the many-body interactions between them. Variants of this approach have been used to parameterize classical force fields and machine learning potentials, applications that benefit from interoperability between quantum chemistry codes. However, there is a dearth of software that provides interoperability yet is purpose-built to handle the combinatorial complexity of fragment-based calculations. To fill this void we introduce “Fragme∩t”, an open-source software application that provides a tool for community validation of fragment-based methods, a platform for developing new approximations, and a framework for analyzing many-body interactions. Fragme∩t includes algorithms for automatic fragment generation and structure modification, and for distance- and energy-based screening of the requisite subsystems. Checkpointing, database management, and parallelization are handled internally and results are archived in a portable database. Interfaces to various quantum chemistry engines are easy to write and exist already for Q-Chem, PySCF, xTB, Orca, CP2K, MRCC, Psi4, NWChem, GAMESS, and MOPAC. Applications reported here demonstrate parallel efficiencies around 96% on more than 1000 processors but also showcase that the code can handle large-scale protein fragmentation using only workstation hardware, all with a codebase that is designed to be usable by non-experts. Fragme∩t conforms to modern software engineering best practices and is built upon well established technologies including Python, SQLite, and Ray. The source code is available under the Apache 2.0 license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

WorkflowHub: a registry for computational workflows

The rising popularity of computational workflows is driven by the need for repetitive and scalable data processing, sharing of processing know-how, and transparent methods. As both combined records of analysis and descriptions of processing steps, workflows should be reproducible, reusable, adaptable, and available. Workflow sharing presents opportunities to reduce unnecessary reinvention, promote reuse, increase access to best practice analyses for non-experts, and increase productivity. In reality, workflows are scattered and difficult to find, in part due to the diversity of available workflow engines and ecosystems, and because workflow sharing is not yet part of research practice. WorkflowHub provides a unified registry for all computational workflows that links to community repositories, and supports both the workflow lifecycle and making workflows findable, accessible, interoperable, and reusable (FAIR). By interoperating with diverse platforms, services, and external registries, WorkflowHub adds value by supporting workflow sharing, explicitly assigning credit, enhancing FAIRness, and promoting workflows as scholarly artefacts. The registry has a global reach, with hundreds of research organisations involved, and more than 800 workflows registered.

97 MATHEMATICS AND COMPUTING↗

Materials Data Science Ontology(MDS-Onto): Unifying Domain Knowledge in Materials and Applied Data Science

Ontologies have gained popularity in the scientific community as a way to standardize terminologies in organizations’ data. Although certain cohorts have created frameworks with rules and guidelines on creating ontologies, there exist significant variations in how Materials Science ontologies are currently developed. We seek to provide guidance in the form of a unified automated framework for developing interoperable and modular ontologies for Materials Data Science that simplifies the ontology terms matching by establishing a semantic bridge up to the Basic Formal Ontology(BFO). This framework provides key recommendations on how ontologies should be positioned within the semantic web, what knowledge representation language is recommended, and where ontologies should be published online to boost their findability and interoperability. Two fundamental components of the MDS-Onto framework are the bilingual package called FAIRmaterials for ontology creation and FAIRLinked, for FAIR data creation. To showcase the practical capabilities of FAIRmaterials, we present two exemplar domain ontologies of MDS-Onto: Synchrotron X-Ray Diffraction and Photovoltaics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Providing a Flexible and Comprehensive Software Stack Via Spack, an Extreme-Scale Scientific Software Stack, and Software Development Kits

To manage the complex demands of modern high-performance computing (HPC), software applications increasingly depend on software developed by other teams, often at other institutions. An HPC software ecosystem approach is required to support dependencies on third-party scientific software. An ecosystem approach provides layers of activity above the individual software product level that promote interoperability, quality improvement, porting, testing, and deployment. The U.S. Exascale Computing Project (ECP) developed its HPC software ecosystem using a three-pronged approach. First, the ECP adopted and invested in Spack, a package manager designed to handle complex HPC package dependencies. Second, the ECP created the Extreme Scale Scientific Software Stack, an effort that supports developing, deploying, and running scientific applications on HPC platforms. Third, the ECP supported software product communities, or software development kits, to develop and promote best practices, improve software interoperability, and other collaborative efforts. This article describes ECP contributions to HPC software ecosystem challenges.

97 MATHEMATICS AND COMPUTING↗

The Vertebrate Breed Ontology: Toward Effective Breed Data Standardization

Abstract Background Limited universally-adopted data standards in veterinary medicine hinder data interoperability and therefore integration and comparison; this ultimately impedes the application of existing information-based tools to support advancement in diagnostics, treatments, and precision medicine. Hypothesis/Objectives A single, coherent, logic-based standard for documenting breed names in health, production, and research-related records will improve data use capabilities in veterinary and comparative medicine. Animals No live animals were used. Methods The Vertebrate Breed Ontology (VBO) was created from breed names and related information compiled from the Food and Agriculture Organization of the United Nations, breed registries, communities, and experts, using manual and computational approaches. Each breed is represented by a VBO term that includes breed information and provenance as metadata. VBO terms are classified using description logic to allow computational applications and Artificial Intelligence–readiness. Results VBO is an open, community-driven ontology representing over 19 500 livestock and companion animal breed concepts covering 49 species. Breeds are classified based on community and expert conventions (e.g., cattle breed) and supported by relations to the breed's genus and species indicated by National Center for Biotechnology Information (NCBI) Taxonomy terms. Relationships between VBO terms (e.g., relating breeds to their foundation stock) provide additional context to support advanced data analytics. VBO term metadata includes synonyms, breed identifiers/codes, and attributed cross-references to other databases. Conclusion and Clinical Importance The adoption of VBO as a standard for breed names in databases and veterinary electronic health records enhances veterinary data interoperability and computability, supporting precision medicine.

Veterinary Sciences↗