Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Ten questions on future and extreme weather data for building simulation and analysis in a changing climate

Weather plays a significant role in building operations as it directly influences HVAC loads and in turn the building energy and thermal performance. In a changing climate, future trends and extreme weather events become critical concerns in the global building decarbonization and clean energy transition. This paper aims to address ten key questions concerning extreme and future weather data for building applications, and more importantly to identify research gaps and guide the curation and selection of future and extreme weather data for use in building performance simulation and assessment. The paper intends to inform architects and engineers, operators, owners, policy makers, and other stakeholders on considering the impacts of future and extreme weather data and adopting strategies for selecting and applying this data in various use cases related to building design, operation, and retrofit for energy efficiency, electrification, and climate resilience.

Yan, Da↗

Tracking Community Building in Open Science

Open Science is enabled by a vibrant community of researchers who regularly engage with the data, from its production to its organization, curation, archiving, dissemination, analysis, and publication. This presentation will examine community building in open science. The NASA Open Science Data Repository (OSDR) makes data available to the public following the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles. OSDR takes open science further with the OS Analysis Working Groups (AWGs) that facilitate community development and promotion. The primary activity of each AWG is to establish and validate analytical processes to generate higher-order data from data housed in OSDR. There are a number of these groups on various topics, including the Animal AWG, Plant AWG, Microbial AWG, Multi-Omics AWG, AI/ML AWG, and the Ames Life Sciences Data Archive (ALSDA) AWG. The international volunteers participating in these AWGs come from academia, citizen science initiatives, industry, and government. They include researchers, principal investigators, professors, trained hobbyists, and students from various domains and disciplines. Anyone may request to join the AWGs, and membership requests are vetted monthly by the group organizers before granting admission. Core to membership is demonstrated expertise through records of training, integrity, work in the professed domain(s), and good community standing. Regular virtual meetings are held for each AWG, with a varying cadence depending on the group's needs and goals. AWG communities share their expertise in research including cutting edge tools, software, frameworks, data formats, and libraries accelerating research collectively. This collaborative approach helps community members cross technology gaps and identify emerging challenges. These diverse communities encompass a wide range of individuals hailing from various sectors within the Science Mission Directorate and beyond. They serve as a means to promote and enhance transparency, accessibility, and inclusion. An annual AWG Symposium brings contributors together in person. Participation in AWGs can be synchronous or asynchronous, with some groups performing most of their work in off hours. Participants gain valuable skills and connections that allow them to add value to their communities and new organizations that they join, resulting in an expanded return on investment for the space life science community. Open science is increasingly a federal mandate and initiatives like NASA's Transform to Open Science and instruments like the Decadal Survey of Biological and Physical Sciences in Space demonstrate the need to carefully consider best practices in this domain. Here, we present greater detail about the makeup and participation metrics of the various AWGs affiliated with OSDR and details of successful peer-reviewed publication campaigns.

Christina M Johnson↗

Tracking Community Building in Open Science

Open Science is enabled by a vibrant community of researchers who regularly engage with the data, from its production to its organization, curation, archiving, dissemination, analysis, and publication. This presentation will examine community building in open science. The NASA Open Science Data Repository (OSDR) makes data available to the public following the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles. OSDR takes open science further with the OS Analysis Working Groups (AWGs) that facilitate community development and promotion. The primary activity of each AWG is to establish and validate analytical processes to generate higher-order data from data housed in OSDR. There are a number of these groups on various topics, including the Animal AWG, Plant AWG, Microbial AWG, Multi-Omics AWG, AI/ML AWG, and the Ames Life Sciences Data Archive (ALSDA) AWG. The international volunteers participating in these AWGs come from academia, citizen science initiatives, industry, and government. They include researchers, principal investigators, professors, trained hobbyists, and students from various domains and disciplines. Anyone may request to join the AWGs, and membership requests are vetted monthly by the group organizers before granting admission. Core to membership is demonstrated expertise through records of training, integrity, work in the professed domain(s), and good community standing. Regular virtual meetings are held for each AWG, with a varying cadence depending on the group's needs and goals. AWG communities share their expertise in research including cutting edge tools, software, frameworks, data formats, and libraries accelerating research collectively. This collaborative approach helps community members cross technology gaps and identify emerging challenges. These diverse communities encompass a wide range of individuals hailing from various sectors within the Science Mission Directorate and beyond. They serve as a means to promote and enhance transparency, accessibility, and inclusion. An annual AWG Symposium brings contributors together in person. Participation in AWGs can be synchronous or asynchronous, with some groups performing most of their work in off hours. Participants gain valuable skills and connections that allow them to add value to their communities and new organizations that they join, resulting in an expanded return on investment for the space life science community. Open science is increasingly a federal mandate and initiatives like NASA's Transform to Open Science and instruments like the Decadal Survey of Biological and Physical Sciences in Space demonstrate the need to carefully consider best practices in this domain. Here, we present greater detail about the makeup and participation metrics of the various AWGs affiliated with OSDR and details of successful peer-reviewed publication campaigns.

Christina M Johnson↗

Open Science for Life in Space: Bioimaging, Data Sharing, and Tools for Knowledge Discovery

Precious space-flown biological experiments have both multi-omic and phenotypic data which NASA strives to make maximally open access for reuse. Currently a number of these space-relevant bioimaging datasets are being reused for AI/ML approaches. NASA Ames Life Science Data Archive and NASA GeneLab are working to make all current and future bioimaging data even more accessible and reusable. Standards for collection and curation are being implemented to enable scientists worldwide access to these data for further discovery and use.

data science↗

Astromaterials Curation Online Resources for Principal Investigators

The Astromaterials Acquisition and Curation office at NASA Johnson Space Center curates all of NASA's extraterrestrial samples, the most extensive set of astromaterials samples available to the research community worldwide. The office allocates ~1500 individual samples to researchers and students each year and has served the planetary research community for 45+ years. The Astromaterials Curation office provides access to its sample data repository and digital resources to support the research needs of sample investigators and to aid in the selection and request of samples for scientific study. These resources can be found on the Astromaterials Acquisition and Curation website at https://curator.jsc.nasa.gov. To better serve our users, we have engaged in several activities to enhance the data available for astromaterials samples, to improve the accessibility and performance of the website, and to address user feedback. We havealso put plans in place for continuing improvements to our existing data products.

Todd, Nancy S.↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

Infusion of AI/ML Technology into Operational NASA Data Systems

NASA has been developing a variety of Artificial Intelligence / Machine Learning technologies related to Earth Observations. In most cases, the full value of such a technology is realized when it is infused into an operational system. NASA’s Earth Science Data Systems program has been formulating repeatable methods to execute technology infusion. These efforts include the Advancing Collaborative Connections for Earth System Science (ACCESS) program, a Technology Infusion Playbook, and an assemblage of working groups investigating methods for infusion collaboration, community development, and capacity building. ESDS has also been executing a pathfinder activity to infuse a machine-learning-driven recommender of science keywords for Earth Observation datasets, which is intended to be used for metadata curation in the Earth Observation System Data and Information System.

C Lynnes↗

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF↗

The NASA Open Science Data Repository: Biomedical Fair Data, Analysis Tools, User Communities, Publications, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

space biology↗

NASA Open Science Data Repository: Biomedical FAIR Data, Analysis Tools, User Communities, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

open access↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter↗

Genesis Solar Wind Interstream, Coronal Hole and Coronal Mass Ejection Samples: Update on Availability and Condition

Recent refinement of analysis of ACE/SWICS data (Advanced Composition Explorer/Solar Wind Ion Composition Spectrometer) and of onboard data for Genesis Discovery Mission of 3 regimes of solar wind at Earth-Sun L1 make it an appropriate time to update the availability and condition of Genesis samples specifically collected in these three regimes and currently curated at Johnson Space Center. ACE/SWICS spacecraft data indicate that solar wind flow types emanating from the interstream regions, from coronal holes and from coronal mass ejections are elementally and isotopically fractionated in different ways from the solar photosphere, and that correction of solar wind values to photosphere values is non-trivial. Returned Genesis solar wind samples captured very different kinds of information about these three regimes than spacecraft data. Samples were collected from 11/30/2001 to 4/1/2004 on the declining phase of solar cycle 23. Meshik, et al is an example of precision attainable. Earlier high precision laboratory analyses of noble gases collected in the interstream, coronal hole and coronal mass ejection regimes speak to degree of fractionation in solar wind formation and models that laboratory data support. The current availability and condition of samples captured on collector plates during interstream slow solar wind, coronal hole high speed solar wind and coronal mass ejections are de-scribed here for potential users of these samples.

Allton, J. H.↗

The Value of Being a Trustworthy Repository

Today, NASA's Earth Observing System Data and Information System (EOSDIS), a system ofactive archives is attaching the CoreTrustSeal to its websites signifying that it merits theconfidence of its user community. But what value does being a trustworthy repository impart to auser? What does it mean to the owners and operators of repositories? What will it mean in thefuture? EOSDIS was started in the 1990s based on a framework of discipline-oriented, geographicallydistributed centers of expertise, named Distributed Active Archive Centers (DAACs). The functionof EOSDIS is to collect Earth Science data sensor measurements (principally those created andneeded by NASA) and manage the data and many derived digital products. EOSDIS providesmany services, including processing, curating, documenting, disseminating, and enabling datadiscovery as well as efficient use of the data. The EOSDIS has been operational over 25 years andmany lessons have been learned relative to the TRUST principles. During the tenure of EOSDIS,many changes have occurred as we have increased the size of the collection from gigabytes totens of petabytes and the distribution of the data to millions of users. We have had severalstages of system evolution that have improved EOSDIS in order to meet both stakeholder andcustomer expectations. This type of evolution is an on-going process to ensure that ourrepositories remain trustworthy. It is also important that our own community of data managersand system engineers add value in being trustworthy. This paper will discuss approaches to change within a large system of Earth Science data and services, while remaining a trustworthyrepository.

Behnke, Jeanne↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Analysis and Review of NASA Earth Science Metadata: How Automation Plays a Role

The Analysis and Review of the Common Metadata Repository (CMR ARC) Team reviews all EOSDIS metadata. The team’s objective is to achieve consistency, correctness, and completeness for all metadata records in the CMR, as well as improve the discoverability of NASA's Earth Science data within the CMR framework. This work is currently being completed at Marshall Space Flight Center. CMR makes a single discovery point possible for NASA's Earth Science data users. The CMR team, in collaboration with three other core metadata teams, contributes to the stewardship of NASA's Earth Science data through a process of continual curation and the ongoing development of the Unified Metadata Model (UMM). A key tool now used in the curation process, referred to as the NASA CMR Dashboard, is an online curation dashboard developed in collaboration with software development company, Element 84. This tool facilitates the review of Earth Science metadata records and subsequent stakeholder collaboration on the resolution of identified issues. A key capability of the new tool is a suite of automated compliance checks written in Python 3.6 that verify the integrity of various metadata elements across multiple standards.

Staton, Patrick↗