Search NASA⌕ Search

SEARCH · Search NASA

Results for “FAIR Data (Findable, Accessible, Interoperable, and Reusable)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Governing Data Findability, Accessibility, Interoperability and Reusability (FAIR) Compliance

The most recent data strategy documents at both the federal and NASA levels stipulate that systems should strive for the data they manage to be Findable, Accessible, Interoperable, and Reusable (FAIR). The NASA Life Sciences Portal (NLSP) has already begun leading efforts in this area for HRP, initiating efforts to comply with the FAIR principles. The broad interpretation of the FAIR principles has led to a plethora of tools that use a splay of metrics specifically but variably developed to judge how compliant data and systems are with the principles. A recent review [3] identified and studied 1,180 metrics across 20 publicly available tools for checking FAIR compliance of data and systems. Because of their very recent development, many organizations and data systems managers and developers have not yet had adequate time or resources to understand these FAIR compliance tools and metrics, their variations in design, accuracy or ease of application to their specific data sets and systems. Thus, it would be best for larger organizations like NASA to approach formulating a strategy for governance of FAIR compliance that can be flexibly applied and is adaptable to an evolving awareness knowledge of FAIR compliance methods and tools. In September 2024, the NASA Science Mission Directorate(SMD) organized a workshop on NASA science data repositories, including the topics of implementing FAIR and governing FAIR compliance across SMD. The initial part of these FAIR discussions focused on developing consensus around required science metadata fields. This is challenging given the diverse nature of NASA’s scientific data portfolio, the variety of metadata models and vocabularies used, and variable level of resources available to curate these data. Later discussion focused on three possible approaches to governing FAIR compliance: distributed, in which various programs, projects or systems define their own methods for assessing FAIR compliance, reporting results up appropriate management lines; centralized, in which higher-level organization(s) specify compliance tools or methods for the various data systems; and multi-level, in which a group comprised of individuals with expertise from multiple levels with organizations is formed to provide guidance and/or specifications for governing FAIR compliance. We report on the recommendations this session yielded, and how these might be shaped specifically to help implement and govern the compliance with FAIR of Human Research Program data and systems.

governance↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching. The use of health countermeasures and biomonitoring systems for space missions are required to counteract space health hazards and to support life to thrive in deep space (e.g., humans, animals, plants, crops; entire ecosystems within spacecrafts/habitats/spacesuits). The development of these mission components will be highly dependent on our understanding of basic biological and health responses to myriad space hazards (ionizing radiation, altered gravitational fields, altered day-night cycles, confined isolation, hostile-closed environments, distance-duration from Earth, planetary dust-regolith, and extreme temperatures/atmospheres). The fast-growing array of space biological and mission telemetry data, which in the past was simply archived after minimal analysis, holds great potential once applied to these mission challenges if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its multi-hierarchical, multi-modal, and heterogenous nature (molecular, cellular, tissue, organ, whole organism, behavior, ecosystem, microbiome; tabular, omics, imaging, video, biospecimen, environmental physical-chemical telemetry). This session focuses on current approaches in this domain such as: making space biological data FAIR (findable, accessible, interoperable, reusable), effective data ingestion/dissemination, observational versus experimental data, Open Science collaborations, data analysis techniques, AI/ML/knowledge graph/modeling methods, and data integration/discovery tools.

open science↗

NASA Open Science Data Repository: Maximizing Spaceflight Bioscience Data

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for data re-analysis and re-use via Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). To address the challenges posed by gaining new knowledge from a vast and diverse amount of biological, health and environmental data in space, the NASA Open Science Data Repository (OSDR - osdr.nasa.gov/bio) plays a crucial role in curating and openly publishing biological data from space-related experiments. Its design incorporates successes and lessons from NASA GeneLab, encompassing not only high-throughput sequencing data but also physiological, phenotypic, and telemetry data. The OSDR makes space biological data FAIR (findable, accessible, interoperable, reusable), and facilitates effective data ingestion, dissemination, and Open Science collaborations. The OSDR also has the capability to integrate human astronaut data with state-of-the-art security and accessibility procedures. We will discuss here several strategies that NASA’s Biological and Physical Science Division have put in place to maximize the return on investment for spaceflight bioscience data.

space biology↗

The Historical Greenland Climate Network (GC-Net) Curated and Augmented Level-1 Dataset

The Greenland Climate Network (GC-Net) consists of 31 automatic weather stations (AWSs) at 30 sites across the Greenland Ice Sheet. The first site was initiated in 1990, and the project has operated almost continuously since 1995 under the leadership of the late Konrad Steffen. The GC-Net AWS measured air temperature, relative humidity, wind speed, atmospheric pressure, downward and reflected shortwave irradiance, net radiation, and ice and firn temperatures. The majority of the GC-Net sites were located in the ice sheet accumulation area (17 AWSs), while 11 AWSs were located in the ablation area, and two sites (three AWSs) were located close to the equilibrium line altitude. Additionally, three AWSs of similar design to the GC-Net AWS were installed by Konrad Steffen's team on the Larsen C ice shelf, Antarctica. After more than 3 decades of operation, the GC-Net AWSs are being decommissioned and replaced by new AWSs operated by the Geological Survey of Denmark and Greenland (GEUS). Therefore, making a reassessment of the historical GC-Net AWS data is necessary. We present a full reprocessing of the historical GC-Net AWS dataset with increased attention to the filtering of erroneous measurements, data correction and derivation of additional variables: continuous surface height, instrument heights, surface albedo, turbulent heat fluxes, and 10 m ice and firn temperatures. This new augmented GC-Net level-1 (L1) AWS dataset is now available at https://doi.org/10.22008/FK2/VVXGUT (Steffen et al., 2023) and will continue to be refined. The processing scripts, latest data and a data user forum are available at https://github.com/GEUS-Glaciology-and-Climate/GC-Net-level-1-data-processing (last access: 30 November 2023). In addition to the AWS data, a comprehensive compilation of valuable metadata is provided: maintenance reports, yearly pictures of the stations and the station positions through time. This unique dataset provides more than 320 station years of high-quality atmospheric data and is available following FAIR (findable, accessible, interoperable, reusable) data and code practices.

Greenland Climate Network↗

TRUST, Trustworthiness and EOSDIS

In recent years there has been considerable attention by the international scientific research and applications community to ensure high quality of data and information management. The terms FAIR (Findable, Accessible, Interoperable, Reusable) data, TRUST (Transparency, Responsibility, User Community, Sustainability, and Technology) principles, and CARE (Collective Benefit, Authority to Control, Responsibility, and Ethics) principles have come into vogue during the last decade. NASA has been managing data and information for over 60 years. NASA’s Earth Observing System Data and Information System (EOSDIS) has been in operation for over 25 years, managing most of NASA’s Earth science data. Trustworthiness is a goal that NASA has always strived to achieve or exceed, because it: enables the success of any NASA science mission; inspires general science research and applications; justifies the cost of operations; contributes to the value of NASA’s Open Data Policy; and influences the long term, historical view for the data collection. Given the recent growth of interest in TRUST principles, it is useful to assess and show how NASA’s attention to trustworthiness maps into those principles. This presentation addresses shows how the various steps that have been taken by the Earth Science Data and Information System (ESDIS) Project in the implementation and evolution of EOSDIS map into the TRUST principles.

Remote Sensing↗

FAIR-ness Assessment of NASA’s Earth Observation System Data and Information System (EOSDIS)

This presentation addresses the challenge of evaluating a multi-disciplinary institutional network of data repositories in operation since 1994 against the relatively recent criteria that constitute FAIR (Findable, Accessible, Interoperable, Reusable) data. NASA’s Earth Observation System Data and Information System (EOSDIS), with its 12 discipline-based Distributed Active Archive Centers (DAACs), preceded the definition and popularization of FAIR by over two decades. An assessment is very useful to describe how well the FAIR principles are met and to identify any improvements needed. In 2020, A “self-assessment” of EOSDIS and DAACs was performed by the ESDIS Project staff and the DAACs from the points of view of human actionability and machine actionability. More recently, a draft of a Science Mission Directorate (SMP) Program Directive (SPD-41a) has been released by NASA Headquarters for comment, where it is recommended that all SMD-funded data should follow the FAIR principles. This presentation is timely to initiate community discussion within the Information Quality Cluster (IQC) of the Earth Science Information Partners (ESIP) and help strategize and develop implementation guidelines for EOSDIS and DAACs to conform to FAIR principles.

Remote sensing↗

NASA’s Safety, Reliability, and Mission Assurance Digital Future

The evolution from “document-centric” to “data-centric” and “model-centric” information leveraging structured data and model-based approaches is at the heart of digital engineering transformational efforts underway across industry and government. It is these approaches that pave the way for data lakes, Authoritative Sources of Truth (ASOTs), and systems- of-systems interoperability and the corresponding transformational benefits thereof. Such benefits include increased data availability, data access equity, data traceability, real-time analytics, batch analytics, and (most importantly) acceleration of the time-to-value and time-to-insights associated with engineering products and analyses. The longer-term benefits of reusability, customization and traceability are even more promising. For Safety and Mission Assurance (SMA), and Mission Success (SMS) activities; realization of such benefits is essential to provide engineers and analysts alike vital information when needed to support critical decision making throughout the entire life cycle. The SMA community often operate in parallel with engineering activities, for which information exchange with relevant context is paramount. Far too often, such information lags key decision points and/or is absent of the robust, integrated, knowledge needed, given inherent barriers associated with traditional document-centric means to data sharing, analysis, and reporting. This paper provides an overview of how NASA’s Office of Safety and Mission Assurance (OSMA) is evolving its policies, standards, guidance, and training to transform to eliminate such barriers, thus realizing the benefits emerging in this new digital era. A roadmap for achieving this digital future is presented along with key building blocks involving use and implementation of concepts such as: Objectives-Hierarchies, Objective-Driven Requirements, Accepted Standards, Safety and Assurance Cases, data digitization (i.e., ontologies, structured data, and model-centric data), FAIR (Findable, Accessible, Interoperable, & Reusable) and/or FAIRUST (Findable, Accessible, Interoperable, Reusable, Understandable, Secure, and Trusted) principles [1]. This paper also describes how OSMA, leveraging the Agency’s overall commitment to Digital Transformation (DT), is using the power of Policy, “Digital” Domain representation, Product Evolution, and Community Outreach and Engagement as part of a strategic vision and roadmap to evolve and transform its SMA organizations to become better able to serve its stakeholders and customers. Future publications will elaborate on these building blocks and deeper concepts.

Authoritative Source of Truth (ASOT),↗

Data Needs to be…

Findable, Accessible, Interoperable, and Reusable (FAIR) data are essential to heliophysics, indeed all scientific research. We make recommendations intended to prioritize resources needed to satisfy FAIR data principles, treating them as a fundamental research infrastructure, rather than a simple research product.

A Halford↗

FAIR Assessments

Presentation on Findable Accessible Interoperable and Reusable (FAIR) Assessment approaches for data.

FAIR↗

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology↗

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism, behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology↗

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez↗

RadLab and the Environmental Data Application Dashboard: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Ionizing radiation in particular has been established in ground-based experiments as being correlated with increased risk of carcinogenesis and cardiovascular and neurological effects. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (osdr.nasa.gov) has developed two Web applications: the Environmental Data Application (EDA) and a radiation-specific RadLab. Each consists of an API (application programming interface) and an associated GUI (graphical user interface) that provide single points of access to the data. To date, OSDR has focused on the sensors from payloads and radiation detectors located on the ISS. The Web applications process telemetry information and associated data, such as spacecraft location and orientation, from multiple international databases. The applications’ request syntax enables users to interrogate these data by craft, sensor type, time range, radiation type (galactic cosmic rays, solar particle events, the contribution of the South Atlantic Anomaly), facilitating arbitrary comparisons of original source data at varying time resolutions. The applications provide programmatic access for use in computational pipelines and GUIs for data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation↗

RadLab: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on multiple spacecraft in and beyond low Earth orbit continuously monitor and collect space radiation data and transmit it back to Earth. These data are of vast importance to space biology research, as ionizing radiation affects living organisms—astronauts and non-human experiment subjects alike—placing them at higher risk of carcinogenesis, degenerative diseases, and radiation sickness. Therefore, knowledge of the biological effects of space radiation is essential for planning future crewed missions beyond low Earth orbit. The RadLab project, initiated by GeneLab and ALSDA (the Open Science Data Repository; OSDR) and sponsored by the NASA Human Research Program, is a new effort aimed at connecting dosimetry data from radiation detectors located on the International Space Station (ISS), as well as other spacecraft. To date, access to these data has been fragmented across space agencies and databases; to address this issue, we have developed an application programming interface (API) and an associated graphical user interface (GUI) designed to provide a single point of access to the data. As of now, OSDR has focused on the detectors located on the ISS, with the long-term goal to establish a self-sustained portal receiving continuous updates through APIs connecting to multiple radiation databases of varying scope, as well as individual investigator contributions. The RadLab API implements a request syntax enabling users to query data by craft, sensor type, timespan, etc, allowing for arbitrary combinations of original source data, thus providing programmatic access for use in computational pipelines, while the GUI facilitates data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation↗

The Environmental Data Application for Analysis of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS and spacecrafts, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (OSDR) has developed a user interface for interrogation of telemetry data: the Environmental Data Application (EDA). The EDA provides the capability to visualize telemetry and radiation data collected on the International Space Station and corresponding ground platforms during the Rodent Research missions. Telemetry data includes temperature, relative humidity, and CO2 levels. Radiation data includes galactic cosmic rays, the contribution of the South Atlantic Anomaly, total radiation dose rate, and accumulated radiation dose. The application allows users to view single missions, compare multiple missions, and view and download summary or full data tables. In summary, the EDA provides GUIs for data visualization and exploration, as well as means for data export, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

telemetry↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Database Design Strategies for Coordinated Simulation and Testing in Additive Manufacturing

The qualification and certification (Q&C) process presents a significant challenge for widespread adoption of additive manufacturing (AM) materials and processes for aerospace applications. A relational database framework will be presented as a tool for data curation of coordinated experimental and computational materials modeling research activities. A comparison of relational and hierarchical data structures in this domain will be emphasized through the evolution of a database design strategy. This framework’s mission is to support the advancement of computational materials-informed Q&C by providing the necessary data infrastructure to trace reliability and reproducibility measures through unified AM materials simulation and experimental testing. FAIR (findable, accessible, interoperable, and reusable) data will be highlighted as a necessary precursor for automation of specific actions, which ultimately reduces the time and expense burden for Q&C. The discussion will be mostly limited to back-end design elements, though a few front-end user experience examples will also be shared.

Qualification↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗