Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The Science Discovery Engine: Connecting Heterogeneous Scientific Data and Information

Transformative science often occurs at the boundaries of different disciplines. Making interdisciplinary science data, software and documentation discoverable and accessible is essential to enabling transformative science. However, connecting this diverse and heterogeneous information is often a challenge due to several factors including the dispersed and sometimes isolated nature of data and the semantic differences between topical areas. NASA’s Science Discovery Engine (SDE) has developed several approaches to tackling these challenges. The SDE is a unified, insightful search experience that enables discovery of NASA’s open science data across five topical areas: astrophysics, biological and physical sciences, Earth science, heliophysics and planetary science. In this presentation, we will discuss our efforts to develop a systematic scientific curation workflow to integrate diverse content into a single search environment. We will also share lessons learned from our work to create a metadata crosswalk across the five disciplines.

Kaylin Bugbee↗

Biomass Harmonization and SAR Analysis with the Multi-mission Algorithm and Analysis Platform (MAAP)

The Multi‐mission Algorithm and Analysis Platform (MAAP) is a collaborative effort between NASA and the European Space Agency (ESA) to support above ground biomass (AGB) research in an open science framework. MAAP brings together relevant data, algorithms, and computing capabilities in a common cloud environment to address the challenges of sharing and processing data from field, airborne and satellite measurements. MAAP was publicly released in October 2021, providing computing capabilities co-located with the data, a collaborative coding and analysis environment, and a set of interoperable tools and algorithms developed to support the estimation and visualization of data. MAAP has allowed scientists from both North America and Europe to collaborate on the generation and analysis/visualization of data derived from multiple, discipline-adjacent missions in an open, collaborative environment that has reached beyond traditional scientific investigation. MAAP has been used to support multiple scientific activities. To date, existing LiDAR data from multiple platforms has been calibrated with field measurements and combined for more comprehensive and accurate estimates of above ground biomass AGB; these LiDAR platforms include airborne (e.g. LVIS), the International Space Station (NASA’s Global Ecosystem Dynamics Investigation (GEDI), and satellites (e.g. ICESat-2). The current challenge is to effectively and seamlessly combine the aforementioned LiDAR-based data with new data sources such as P-band RADAR from ESA’s upcoming BIOMASS mission, existing ESA Sentinel-1 C-band SAR, and the 30 PB/yr of high cadence global coverage L-band SAR data from the upcoming NASA-ISRO SAR (NISAR) mission. Recent analysis using MAAP merged ICESat-2 and optical data (Harmonized Landsat Sentinel) produced the most comprehensively precise estimate of boreal-wide AGB to date. Another effort using MAAP is the production and open distribution of global comparisons of AGB map estimates, including from ICESat-2 and GEDI, to bolster stakeholder uptake for policy applications. These map estimates will feed into the Intergovernmental Panel on Climate Change (IPCC) database, likely aiding the next Global Carbon Stocktake of the UNFCCC. Furthermore, the biomass retrieval intercomparison exercise BRIX-2 could benefit from the MAAP providing standardized test cases (based on airborne campaign and spaceborne data) allowing the community to develop and apply retrieval algorithms based on these test cases, while forthcoming SAR data training curricula could also use the MAAP as a teaching and learning platform. The MAAP is meeting the challenges inherent in international, open science collaboration and large scale computing with a platform that is entirely open source and cloud native, using open standards for data access, manipulation, protocols, and formats. The MAAP data system consists of a dedicated data store whose data is indexed in an online catalog conforming to established metadata, application programmatic interfaces (APIs), and service interface standards, using an implementation of the open sourced NASA Common Metadata Repository. Federation of user identities allows users from either NASA or ESA to access and consume services from the other using a unified metadata catalog for the data utilized across the ESA and NASA MAAP platforms. Similarly, we are exploring how to increase interoperability to achieve a common approach to packaging, orchestrating and executing algorithms, with interoperable access to data for subsetting, fast browse, and cloud-optimized access, all using interoperable standards such as those from the Open Geospatial Consortium (OGC). Designed for interoperability, ESA and NASA utilize a common architecture for the software platform. It provides a cloud-based algorithm development environment (ADE) that enables scientists to develop algorithms collaboratively with access to the MAAP data catalog as well as other data archives. MAAP provides an Eclipse Che-based ADE supporting both Python and R languages, popular in this biomass community. Algorithms developed and containerized within the ADE can be deployed to run to thousands of computational nodes in the MAAP’s data processing system (DPS), dramatically speeding up processing and giving scientists a rapid, iterative turnaround of results. NASA’s implementation of the DPS is based on the Hybrid Science Data System (HySDS) framework, used by NASA flight projects to produce Earth science standard products.

cloud computing↗

Geocam Space: Enhancing Handheld Digital Camera Imagery from the International Space Station for Research and Applications

Handheld astronaut photography of the Earth has been collected from the International Space Station (ISS) since 2000, making it the most temporally extensive remotely sensed dataset from this unique Low Earth orbital platform. Exclusive use of digital handheld cameras to perform Earth observations from the ISS began in 2004. Nadir viewing imagery is constrained by the inclined equatorial orbit of the ISS to between 51.6 degrees North and South latitude, however numerous oblique images of land surfaces above these latitudes are included in the dataset. While unmodified commercial off-the-shelf digital cameras provide only visible wavelength, three-band spectral information of limited quality current cameras used with long (400+ mm) lenses can obtain high quality spatial information approaching 2 meters/ground pixel resolution. The dataset is freely available online at the Gateway to Astronaut Photography of Earth site (http://eol.jsc.nasa.gov), and now comprises over 2 million images. Despite this extensive image catalog, use of the data for scientific research, disaster response, commercial applications and visualizations is minimal in comparison to other data collected from free-flying satellite platforms such as Landsat, Worldview, etc. This is due primarily to the lack of fully-georeferenced data products - while current digital cameras typically have integrated GPS, this does not function in the Low Earth Orbit environment. The Earth Science and Remote Sensing (ESRS) Unit at NASA Johnson Space Center provides training in Earth Science topics to ISS crews, performs daily operations and Earth observation target delivery to crews through the Crew Earth Observations (CEO) Facility on board ISS, and also catalogs digital handheld imagery acquired from orbit by manually adding descriptive metadata and determining an image geographic centerpoint using visual feature matching with other georeferenced data, e.g. Landsat, Google Earth, etc. The lack of full geolocation information native to the data makes it difficult to integrate astronaut photographs with other georeferenced data to facilitate quantitative analysis such as urban land cover/land use classification, change detection, or geologic mapping. The manual determination of image centerpoints is both time and labor-intensive, leading to delays in releasing geolocated and cataloged data to the public, such as the timely use of data for disaster response. The GeoCam Space project was funded by the ISS Program in 2015 to develop an on-orbit hardware and ground-based software system for increasing the efficiency of geolocating astronaut photographs from the ISS (Fig. 1). The Intelligent Robotics Group at NASA Ames Research Center leads the development of both the ground and on-orbit systems in collaboration with the ESRS Unit. The hardware component consists of modified smartphone elements including cameras, central processing unit, wireless Ethernet, and an inertial measurement unit (gyroscopes/accelerometers/magnetometers) reconfigured into a compact unit that attaches to the base of the current Nikon D4 camera - and its replacement, the Nikon D5 - and connects using the standard Nikon peripheral connector or USB port. This provides secondary, side and downward facing cameras perpendicular to the primary camera pointing direction. The secondary cameras observe calibration targets with known internal X, Y, and Z position affixed to the interior of the ISS to determine the camera pose corresponding to each image frame. This information is recorded by the GeoCam Space unit and indexed for correlation to the camera time recorded for each image frame. Data - image, EXIF header, and camera pose information - is transmitted to the ground software system (GeoRef) using the established Ku-band USOS downlink system. Following integration on the ground, the camera pose information provides an initial geolocation estimate for the individual film frame. This new capability represents a significant advance in geolocation from the manual feature-matching approach for both nadir and off-nadir viewing imagery. With the initial geolocation estimate, full georeferencing of an image is completed using the rapid tie-pointing interface in GeoRef, and the resulting data is added to the Gateway to Astronaut Photography of Earth online database in both Geotiff and Keyhole Markup Language (kml) formats. The integration of the GeoRef software component of Geocam Space into the CEO image cataloging workflow is complete, and disaster response imagery acquired by the ISS crew is now fully georeferenced as a standard data product. The on-orbit hardware component (GeoSens) is in final prototyping phase, and is on-schedule for launch to the ISS in late 2016. Installation and routine use of the Geocam Space system for handheld digital camera photography from the ISS is expected to significantly improve the usefulness of this unique dataset for a variety of public- and private-sector applications.

Stefanov, William L.↗

Publishing Variables Archived at GES DISC to Earth System Grid Federation (ESGF)

We present a straightforward and low-cost approach to publish variables archived at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) to the Earth System Grid Federation (ESGF). An ESGF publication requires a single standard-name variable aggregated over time to facilitate data inter-comparison. It also contains significant metadata to enable searching in ESGF. We look up standard names on high demand in ESGF search history, and using OPeNDAP and NcML technologies we aggregate the corresponding variables available in the GES DISC archive with augmented metadata required by CMIP6 and obs4MIPs Data Specification version 2.1. At this writing 10 variables from a standard product of the Atmospheric Infrared Sounder along with the Tech Notes are published in ESGF by NASA Center for Climate Simulation (NCCS). Users can view, analyze, and subset remotely, and download these aggregated variables via links in any ESGF node after searching. We plan to work on and publish more variables and data from different NASA missions and experiments in our archive.

Fan Fang↗

Enabling Cloud Services and Enhanced Data Discovery With Earthdata-Varinfo

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. Each of these collections can contain hundreds of variables that describe each measurement within the dataset, therefore an automated method for generating UMM-Var records is necessary. The Unified Metadata Model for Variables (UMM-Var) provides a framework for variable metadata records in NASA’s Common Metadata Repository (CMR). The Python tool, earthdata-varinfo, was developed to solve this problem of automating the curation of UMM-Var records. Given either a collection DMR file or a netCDF-4 file, earthdata-varinfo can scrape variable metadata and return a CMR compliant UMM-Var record. Earthdata-varinfo can generate thousands of UMM-Var records in a matter of seconds, thus enabling subsetting capabilities and enhancing data discovery.

Eni Awowale↗

Global Data Assembly Center (GDAC) Report to the GHRSST Science Team

In 2012-2013 the Global Data Assembly Center (GDAC) at NASA's Physical Oceanography Distributed Active Archive Center (PO.DAAC) continued its role as the primary clearinghouse and access node for operational GHRSST data streams, as well as its collaborative role with the NOAA Long Term Stewardship and Reanalysis Facility (LTSRF) for archiving. Our presentation reported on our data management activities and infrastructure improvements since the last science team meeting in 2012.

user reports↗

Datacasting V3.0

Datacasting V3.0 provides an RSSbased feed mechanism for publishing the availability of Earth science data records in real time. It also provides a utility for subscribing to these feeds and sifting through all the items in an automatic manner to identify and download the data records that are required for a specific application. Datacasting is a method by which multiple data providers can publish the availability of new Earth science data and users download those files that meet a predefined need; for example, to only download data files related to a specific earthquake or region on the globe. Datacasting is a server-client architecture. The server-side software is used by data providers to create and publish the metadata about recently available data according to the Datacasting RSS (Really Simple Syndication) specification. The client software subscribes to the Datacasting RSS and other RSS-based feeds. By configuring filters associated with feeds, data consumers can use the client to identify and automatically download files that meet a specific need. On the client side, a Datacasting feed reader monitors the server for new feeds. The feed reader will be tuned by the user, via a graphical user interface (GUI), to examine the content of the feeds and initiate a data pull after some criteria are satisfied. The criteria might be, for example, to download sea surface temperature data for a particular region that has cloud cover less than 50% and during daylight hours. After the granule is downloaded to the client, the user will have the ability to visualize the data in the GUI. Based on the popular concept of podcasting, which gives listeners the capability to download only those MP3 files that match their preference, Earth science Datacasting will give users a method to download only the Earth science data files that are required for a particular application.

Bingham, Andrew W.↗

NASA GeneLab: Open Science for Life in Space

NASA’s GeneLab helps scientists understand how the fundamental building blocks of life – DNA, RNA, proteins, and metabolites – change from exposure to the space environment including microgravity and cosmic radiation exposure. GeneLab does so by providing fully coordinated epigenomics, genomics, transcriptomics, proteomics, and metabolomics data (collectively known as omics data) alongside essential metadata describing each spaceflight and space-relevant experiment. The open-access GeneLab repository currently consists of over 300 omics datasets generated by biological experiments, involving various model organisms, that are relevant to spaceflight. In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab has started processing and analyzing these datasets to generate differential gene expression data and identify biological and physiological pathways that are dysregulated as a result of spaceflight. To aide GeneLab’s efforts to harmonize and democratize space-relevant omics data, over 130 scientists have joined one of four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG) and together helped develop and adopted standard data analysis workflows for all data types available in GeneLab. Currently, the GeneLab Data System includes a data repository with federated search capability, an online controlled-access toolshed powered by "Galaxy" for users to process data with vetted standard workflows, a workspace for data sharing, a data submission portal, and the ability to browse and visualize transcriptomics processed data. The user interface was designed to be accessible to a broad variety of users, including high school and college students who can use it to learn about omics data analysis and space biology. The visualization portal enhances GeneLab’s ability to democratize omics data by removing the need for bioinformatics expertise to interpret transcriptomics data hosted on GeneLab. This presentation will provide an over-view of NASA’s GeneLab including how to navigate the GeneLab Data System and will conclude by providing resources for opportunities to work with GeneLab and NASA at large.

Amanda M Saravia-Butler↗

Laying the Foundations for FAIR-er Science: ISA and the LSDA Data Submission Process in NASA's Evolving Data Management Environment

The Life Sciences Data Archive (LSDA) archives data resulting from research on the effects of spaceflight on humans and the development of countermeasures to mitigate spaceflight hazards. Archivists work with researchers to ensure that unique and high value data products and their metadata are preserved and managed to support current and future research. Currently, LSDA is updating its procedures and data submission requirements in response to the evolving data preservation environment at NASA. LSDA is implementing best practices for research data management through the establishment of clear data submission guidelines, integration of the FAIR (Findability, Accessibility, Interoperability, Reusability) principles, and use of the ISA (Investigation, Study, Assay) research metadata framework for data discoverability and transparency into the data management processes. These changes directly impact LSDA’s requirements for research data submissions. The newly revised Research Data Submission Agreement (RDSA), formerly the Data Submission Agreement (DSA), introduces ISA-compatible metadata collection standards to LSDA’s process. Adherence to LSDA’s data submission guidelines enhances the FAIR-ness of the repository’s collections for future users. This presentation will discuss (1) how submission of research data and associated metadata are impacted by current data management policies, (2) benefits of the adoption of FAIR principles and the ISA metadata framework for retrospective studies utilizing existing LSDA datasets and historic data collections, and (3) the support LSDA will provide to researchers during this transition.

LSDA↗

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation↗

Data Integrity Challenges in NASA Giovanni

The Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni) is an online tool developed by the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), one of 12 NASA Science Mission Directorate Data Centers (DAACs) to analyze and visualize NASA remote sensing and model data without downloading data and software. As of this writing, over 2000 Earth satellite and model variables are available in Giovanni, including several well-known NASA satellite missions (e.g., TRMM, GPM) and projects (e.g., MERRA-2, GPCP). There are twenty-two plots provided by Giovanni that can be used to analyze, compare, and explore Earth data across disciplines. Results can be shared with colleagues and downloaded for further analysis. Giovanni has helped publish over 3000 referral papers over the years. As open science policies roll in, data integrity has become a major challenge for Giovanni and other tools. For integrity, both data and workflows must be transparent. FAIR-compliant data, including input, intermediate, and result products, as well as their associated statistics, metadata, and information, are needed. The NASA Data Product Development Guide for Data Producers provides a key resource on how to develop FAIR-compliant data products. Data quality information is also needed from data producers and analysis services like Giovanni. The workflow part is quite challenging and requires workflow management improvements, such as recording workflows and making them available to users. In this presentation, we will discuss the data integrity challenges in Giovanni.

data analysis, visualization↗

ncompare: A Python Package for Comparing netCDF Structures

Earth science researchers and data engineers have a common problem: they often need to compare data files to see what is different between them. A lot of time is spent developing code to test differences. When it comes to comparing multidimensional data file formats like netCDFs (Network Common Data Form), this is particularly challenging and time-consuming, since there is frequently a need to evaluate the differences between dimension sizes, variable structures, and variable attributes, especially for regression testing. Since netCDFs are widely used in Earth science — with climate models, oceanographic or atmospheric reanalyses, and observational data — improved means of evaluating netCDF files can help enable a wide range of applications. We have developed a reusable open source approach through `ncompare`, which is a Python package for comparing netCDF structures [[https://github.com/nasa/ncompare]]. The `ncompare` tool compares the structure of two Network Common Data Form (NetCDF) files at the command line. It facilitates rapid comparisons by generating a formatted display of the matching and non-matching groups, variables, and associated metadata between two NetCDF datasets. The user has the option to colorize the terminal output for ease of viewing, and `ncompare` can optionally save comparison reports in text, comma-separated value (CSV), and/or Microsoft Excel formats. Despite the availability of tools (such as ncmpidiff or nccmp) that compare the values of variables, there was not previously a readily available, Python-based tool for rapid visual comparisons of group and variable structures, attributes, and chunking. `ncompare` was developed at NASA’s Atmospheric Science Data Center (ASDC) and is a collaboration with NASA Openscapes [[https://nasa-openscapes.github.io]] mentors across 11 of NASA’s data centers. Openscapes’ overarching vision is to support scientific researchers using NASA Earthdata as they migrate their workflows to the cloud. Relevant links: - https://github.com/nasa/ncompare - https://github.com/pyOpenSci/software-submission/issues/146 - https://nasa-openscapes.github.io

Daniel Kaufman↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods. REFERENCES [1] Open science in space. Nature Medicine, 2021. 27(9): p. 1485-1485. [2] Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. [3] Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. [4] Whetzel, P.L., et al., BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res, 2011. 39(Web Server issue): p. W541-5.

informatics↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods.

knowledge↗

Building access and community standards for opacity data at the onset of next-generation atmosphere observations

The characterization of a diverse set of exoplanet atmosphere observations, ranging from hot gas giants to small temperate rocky worlds, will be one of the legacies of upcoming facilities such as the James Webb Space Telescope (JWST). Our understanding and interpretation of such observations will hinge on our ability to link observations with atmospheric theoretical studies that critically rely on fundamental molecular and atomic opacities. Computing such opacities is a highly non-trivial and inaccessible process which requires several terabytes of available disk space, hours of CPU time per pressure-temperature combination, and requires users to carefully aggregate line lists data from various sources, which limits access and intercomparison of opacity data in the exoplanet community. Here we present MAESTRO (Molecules and Atoms in Exoplanet Science: Tools and Resources for Opacities) an opacity database that can be accessed by the community via a web interface and python API. MAESTRO was built with community input to create a version-controlled opacity database that is easily queryable, includes informative metadata to ensure reproducibility, and exports relevant citations for inclusion in publications. Scheduled for community release in 2022, MAESTRO will prove to be an invaluable community resource in the era of JWST and beyond.

Natasha Batalha↗

Value-added Data Services at the Goddard Earth Sciences Data and Information Services Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), in addition to serving the Earth Science community as one of the major Distributed Active Archives Centers (DAACs), provides much more than just data. Among the value-added services available to general users are subsetting data spatially and/or by parameter, online analysis (to avoid downloading unnecessarily all the data), and assistance in obtaining data from other centers. Services available to data producers and high-volume users include consulting on building new products with standard formats and metadata and construction of data management systems. A particularly useful service is data processing at the DISC (i.e., close to the input data) with the users algorithm. This can take a number of different forms: as a configuration-managed algorithm within the main processing stream; as a stand-alone program next to the on-line data storage; as build-it-yourself code within the Near-Archive Data Mining (NADM) system; or as an on-the-fly analysis with simple algorithms embedded into the web-based tools. Partnerships between the GES DISC and scientists, both producers and users, allow the scientists to concentrate on science, while the GES DISC handles the data management, e.g., formats, integration, and data processing. The existing data management infrastructure at the GES DISC supports a wide spectrum of options: from simple data support to sophisticated on-line analysis tools, producing economies of scale and rapid time-to-deploy. At the same time, such partnerships allow the GES DISC to serve the user community more efficiently and to better prioritize on-line holdings. Several examples of successful partnerships are described in the presentation.

Leptoukh, Gregory G.↗

TRMM Version 7 Near-Realtime Data Products

The TRMM data system has been providing near-realtime data products to the community since late 1999. While the TRMM project never had near-realtime production requirements, the science and applications communities had a great interest in receiving TRMM data as quickly as possible. As a result these NRT data are provided under a best-effort scenario but with the objective of having the swath data products available within three hours of data collection 90% of the time. In July of 2011 the Joint Precipitation Measurement Missions Science Team (JPST) authorized the reprocessing of TRMM mission data using the new version 7 algorithms. The reprocessing of the 14+ years of the mission was concluded within 30 days. Version 7 algorithms had substantial changes in the data product file formats both for data and metadata. In addition, the algorithms themselves had major modifications and improvements. The general approach to versioning up the NRT is to wait for the regular production algorithms to have run for a while and shake out any issues that might arise from the new version before updating the NRT products. Because of the substantial changes in data/metadata formats as well as the algorithm improvements themselves, the update of NRT to V7 followed an even more conservative path than usual. This was done to ensure that applications agencies and other users of the TRMM NRT would not be faces with short-timeframes for conversion to the new format. This paper will describe the process by which the TRMM NRT was updated to V7 and the V7 data products themselves.

Tocker, Erich Franz↗

The NASA Open Science Data Repository: Biomedical Fair Data, Analysis Tools, User Communities, Publications, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

space biology↗