Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

Summary of Responses to the Request for Information (RFI) on Partnerships for Transformational Artificial Intelligence Models

The Department of Energy (DOE) issued a Request for Information (RFI) in December 2025 inviting public comments regarding partnerships for transformational Artificial Intelligence (AI) models for the Genesis Mission Consortium, a public-private partnership platform. This RFI solicited feedback from industry, nonprofit organizations, universities, independent research organizations and other stakeholders. Specifically, the RFI asked three questions on (1) mobilizing DOE National Laboratories to curate the scientific data in a responsible and privacy-preserving manner, (2) the extent to which existing general-purpose AI models can be leveraged and which scientific disciplines are priorities for such model development, and (3) mechanisms by which these AI models can be provided to scientific communities. This document summarizes the input from 194 unique nonproprietary responses from businesses, universities, nonprofit organizations, research institutes and laboratories as well as a variety of other contributors, including individual contributions.

97 MATHEMATICS AND COMPUTING↗

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine↗

Objective Structured Clinical Evaluation (OSCE) of an Artificial Intelligence (AI) Clinical Decision Support System (CDSS) Tool

BACKGROUND Objective Structured Clinical Evaluations (OSCEs) have long been established as a robust methodology for summative assessment of clinical skills and decision-making during medical education. The recent integration of Artificial Intelligence (AI) into clinical decision-making processes has prompted the need for novel evaluation frameworks to assess the efficacy and reliability of AI clinical decision support system (CDSS) tools. This abstract outlines the process of quantitatively evaluating a novel CDSS (“Doc in a Box” Google 2024) trained on curated medical spaceflight data in the psychomotor domain as it interfaces with a human volunteer acting as the crew medical officer (CMO). PURPOSE The AI CDSS under review was developed as part of the Lunar Command and Control Interoperability (LuCCI) project, which is intended to address a gap in how Lunar Surface Systems (LSS) would interoperate across multiple programs, commercial partners, and international partners. The project objective is to define, prototype, integrate, and evaluate an interoperable lunar command, control, data, and software reference architecture to enable autonomy and informatics capability through common standards across LSS. A multi-modal AI-based CDSS compatible with Federated LSS will assist clinicians in diagnosing and managing complex medical conditions by providing evidence-based recommendations through predictive analytics. Given the critical role of decision-support as NASA continues to evolve its Earth-independent medical operations (EIMO), it is imperative to ensure that such AI tools perform reliably and align with clinical standards during progressive lunar and Martian exploration class missions. METHODS The OSCE framework, traditionally used for evaluating human clinicians, was adapted to assess the AI tool's decision-making capabilities in simulated clinical scenarios. In this adapted OSCE, the AI CDSS was tested across a series of structured clinical scenarios designed to mimic real-life spaceflight patient cases. These scenarios included a range of conditions and complexities, allowing for comprehensive assessment of the tool's performance. Key evaluation metrics included accuracy of diagnosis, timeliness of decision-making, and appropriate recommendations for therapies. The OSCE was scored by human physician evaluators who assessed the AI's recommendations in comparison with expert clinicians' medical decision making to ensure alignment with best practices and the standard of care. RESULTS Preliminary results indicate that the AI CDSS demonstrated high accuracy in diagnostic recommendations and decision support across various scenarios. However, certain limitations were noted, such as occasional discrepancies in handling complex or nuanced cases that required a more contextual understanding. Additionally, the tool scored higher on the diagnostic portion of the rubric, with lower scores in the therapeutic recommendations. These findings highlight the importance of continuous refinement and validation of AI tools through rigorous evaluation frameworks like the OSCE. The adaptation of OSCEs for AI tools presents several advantages, including a structured and reproducible approach to evaluation, the ability to test AI systems in diverse clinical scenarios, and the opportunity to benchmark AI performance against established clinical standards to permit charting of future progress as aerospace medicine evolves as a discipline. Remaining challenges include ensuring that these evaluations capture the full spectrum of clinical decision-making scenarios that will be confronted by CMOs during missions and adequately reflecting real-world variability of the austere spaceflight environment. CONCLUSION Employing OSCEs to evaluate AI clinical decision support tools offers a promising approach to validating their clinical utility and efficacy. This methodology not only provides insights into the tool's performance but also fosters ongoing improvement and alignment with standard of care practices. Future research should focus on refining these evaluation processes and addressing limitations to enhance the integration of AI tools in clinical spaceflight settings. REFERENCES Scott S, Hearns V, Barker MA. Testing Clinical Skills: A Look at the OSCE and USMLE Clinical Skills Exams. S D Med. 2019 Oct;72(10):451-453. Majumder MAA, Kumar A, Krishnamurthy K, Ojeh N, Adams OP, Sa B. An evaluative study of objective structured clinical examination (OSCE): students and examiners perspectives. Adv Med Educ Pract. 2019 Jun 5;10:387-397. Karam VY, Park YS, Tekian A, Youssef N. Evaluating the validity evidence of an OSCE: results from a new medical school. BMC Med Educ. 2018 Dec 20;18(1):313.

Ariana M Nelson↗

Thermal Performance of Spandrel Assemblies in Glazed Wall Systems: Laboratory Test Design – Challenges and Test Results

Accurate thermal performance calculation procedures for opaque spandrel areas in curtain wall and window wall systems are essential for rating systems when comparing spandrel systems. However, there is a lack of consensus in thermal modeling needed for accurately characterizing heat transfer through spandrel assemblies due to the complex arrangement of materials and structural components. Several studies indicate that conventional 2D thermal simulations may overestimate R-values by 30% compared to physical testing and 3D simulations. Detailed simulations and well-curated laboratory test data are necessary to build confidence in simulation models, which will later be used to develop correlations to improve widely used conventional 2D thermal simulations. This study aims to experimentally test heat transfer through various spandrel assemblies to validate 3D simulation models. Also, the challenges of conducting a thorough testing design along with the solutions would be documented. The team developed a design for testing spandrel assemblies, making appropriate modifications to the existing heat, air, and moisture (HAM) chamber to accommodate the testing needs. Two moveable baffles were designed and fabricated to guide airflow direction parallel to the test article surface. The data acquisition capabilities in the chamber were upgraded to add more than two hundred sensors to the climate and indoor side of the chamber. The goal is to provide a quality dataset for validating complex 3D modeling simulations, which will be used to develop improved thermal simulation techniques that more accurately represent the thermal behavior of spandrel assemblies and their integration within the building envelope. This paper will summarize the results for the boundary conditions of the testing and the temperature variation across different locations of the spandrel assemblies.

Kunwar, Niraj [ORNL] (ORCID:0000000263457652)↗

Compilation of Experimental Yield Data for Spontaneous Fission of 252 Cf

We present a comprehensive compilation and curation of experimental fission yield (FY) data for the spontaneous fission of 252 Cf, extracted from the EXFOR database. The compilation follows a structured methodology developed for prior compilations of neutron-induced fission yields, and incorporates both independent (IFY) and cumulative (CFY) yields. A total of 62 datasets were reviewed, with entries spanning from 1955 to 2021. A significant portion of the literature reports pre-neutron emission yields, which were excluded from the present compilation due to limitations in format compatibility. Each accepted dataset was processed into a standardized JSON format, including metadata, uncertainties, and bibliographic references. Where available, decay radiation information was used to update the FY data using the latest ENSDF evaluations; 237 data points were corrected accordingly. These corrections are fully traceable and preserve original values. The result is a curated dataset suitable for use in nuclear data evaluations. This work is part of an ongoing effort to modernize the handling of FY data and provide evaluators with high-quality, machine-readable experimental inputs

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Blast from the Past: ASDC Curation for NASA Suborbital Legacy Missions to Promote Data Discovery and Accessibility

NASA has an extensive history of conducting suborbital field campaigns to further advances in atmospheric sciences. Beginning with the Chemical Instrument Test and Evaluation (CITE) conducted in 1983-1984, NASA has completed many suborbital campaigns over the past three decades. Since the early 2010s, suborbital missions are typically assigned to a NASA Distributed Active Archive Center (DAAC) prior to the mission for long-term archival and distribution. Efforts are being made by NASA’s Earth Science Data and Information System (ESDIS) Project and the Airborne Data Management Group (ADMG) to assign legacy missions to DAACs for permanent archival and distribution, so that these valuable datasets remain to be available to the scientific community. NASA’s Atmospheric Science Data Center (ASDC) has been named the assigned DAAC for nearly 20 atmospheric composition legacy missions, including missions conducted as part of the Global Tropospheric Experiment (GTE) and expects to be named the assigned DAAC for more of these missions over the next few years. The primary goal of the ASDC is to provide access to the datasets as they are currently formatted to the broad user community and enhance their findability and accessibility. However, data reporting standards have evolved significantly since 1983 and the datasets span a wide variety of file formats, including text, Ames, GTE, and ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), and the amount of metadata and relevant information included in the files also varies greatly and can not be readily extracted without subject matter knowledge. This has caused challenges for the ASDC’s suborbital metadata extraction pipeline in ensuring that accurate and necessary metadata is being provided for the missions by all the ASDC’s existing search mechanisms. To make the data more findable and accessible, the ASDC has begun researching ways to further enhance the datasets, including distributing value-added products (i.e. consistent file format such as ICARTT or netCDF), adding standard names from the ESDIS Standards Coordination Office (ESCO)-approved Atmospheric Composition Variable Standard Names Convention (ACVSNC), and creating outreach materials such as ArcGIS StoryMaps, User Guides, and Micro Articles, providing overviews of the missions and what type of data was collected during the missions. These efforts also help support NASA’s Open-Source Science by enhancing the FAIRness of the legacy data products. This presentation will review the ASDC’s ongoing efforts, progress made, and future plans for legacy missions.

Megan Buzanowicz↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗

Citizen Science Approach for Searching and Curating Literature of the Effects of Spaceflight on Cardiovascular Outcomes in Rodents and Humans

The spaceflight environment causes significant changes to the structure and function of the cardiovascular system, including fluid redistribution, alterations in blood pressure, and changes in cardiac output. The goal of this project is to quantitatively summarize the data on the effects of actual or simulated microgravity and radiation exposure resulting from spaceflight on the cardiovascular system. As the first step, a group of investigators approached through a collaboration of the Ames Life Science Data Archive (ALSDA) Analysis Working Group developed a list of relevant cardiovascular search terms. Based on these, medical librarians generated and executed the search strategy in Medline, CINAHL, Embase and NASA repositories. In parallel, we recruited students and young professionals from various space industry-affiliated organizations, resulting in ~100 individuals joining. With this program we aimed to reach students and young people underrepresented in STEM, including first-generation, female, minorities, disadvantaged backgrounds, fostered individuals, etc. These individuals completed a virtual training course on the nature and methodologies of the project. Following this, the participants were structured into teams with more senior/experienced individuals designated as team leaders. Currently, the teams are screening approximately 15,000 studies using the systematic review tool, Covidence. Teams will be extracting and curating data for meta-analysis of the cardiovascular spaceflight literature, but also extracting, submitting, and curating appropriate datasets into the new ALSDA submission portal and repository. This effort will result in collaborative publications based upon the literature meta-analyses, and a number of publicly accessible datasets for reuse, modeling, machine learning, and knowledge graph-type approaches. This approach reduces the length of time to complete title/abstract screening time from 1-2 years needed for this volume of studies, to 3-4 months, while also providing a unique, open-access educational experience to space research and training in knowledge synthesis tools to interested individuals.

space biology↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

SG50 Data-format Specifications Document for the Automatically Readable, Comprehensive, and Curated Experimental Reaction Database MEDUSAL

The aim of this document is to lay out a first draft of the specifications for the MEDUSAL database (Machine-readable Experimental Data User App & Library) that is being described by OECD/NEA/WPEC SG-50. The EXFOR database (Otuka et al., 2014) has a format that is based on code-value pairs, and a significant portion of the information in the EXFOR entry is contained in free text sections. Several high-level requirements for the MEDUSAL database, as laid out in the Use Cases and Requirements Working Paper (citation), relate to the definition of the specifications

Nuclear Criticality Safety Program (NCSP)↗

Advances in Small Particle Handling of Astromaterials in Preparation for OSIRIS-REx and Hayabusa2: Initial Developments

The Astromaterials Acquisition and Curation office at NASA Johnson Space Center has established an Advanced Curation program that is tasked with developing procedures, technologies, and data sets necessary for the curation of future astromaterials collections as envisioned by NASA exploration goals. One particular objective of the Advanced Curation program is the development of new methods for the collection, storage, handling and characterization of small (less than 100 micrometer) particles. Astromaterials Curation currently maintains four small particle collections: Cosmic Dust that has been collected in Earth's stratosphere by ER2 and WB-57 aircraft, Comet 81P/Wild 2 dust returned by NASA's Stardust spacecraft, interstellar dust that was returned by Stardust, and asteroid Itokawa particles that were returned by the JAXA's Hayabusa spacecraft. NASA Curation is currently preparing for the anticipated return of two new astromaterials collections - asteroid Ryugu regolith to be collected by Hayabusa2 spacecraft in 2021 (samples will be provided by JAXA as part of an international agreement), and asteroid Bennu regolith to be collected by the OSIRIS-REx spacecraft and returned in 2023. A substantial portion of these returned samples are expected to consist of small particle components, and mission requirements necessitate the development of new processing tools and methods in order to maximize the scientific yield from these valuable acquisitions. Here we describe initial progress towards the development of applicable sample handling methods for the successful curation of future small particle collections.

Snead, C. J.↗

Expanding Repository Data Available For Sharing and Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

Biology↗

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

The Unified Phenotype Ontology : a framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

59 BASIC BIOLOGICAL SCIENCES↗

Electricity Baseline 2022

The Electricity Baseline (2022) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2022" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569193. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; data inventory↗