Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Sentinel-5P/TROPOMI and S-NPP/OMPS Data Support at GES DISC

The TROPspheric Monitoring Instrument (TROPOMI) on the Sentinel-5 Precursor (Sentinel-5P) is the first of the Atmospheric Composition Sentinels by the European Space Agency (ESA) that provides measurements of ozone, NO2, SO2, CH4, CO, formaldehyde, aerosols and cloud at high spatial, temporal and spectral resolutions. The early afternoon orbit of Sentinel-5P mission provides a strong synergy with the U.S. Suomi National Polar-orbiting Partnership (S-NPP) satellite, especially in that the S-NPP Ozone Monitoring and Profiling Suite (OMPS) facilitates high vertically resolved stratospheric and lower mesospheric ozone profiles. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) supports over a thousand data collections in the Focus Areas of Atmospheric Composition, Water & Energy Cycles, and Climate Variability and it is the Distributed Active Archive Center (DAAC) that is curating both offline Sentinel-5P TROPOMI and S-NPP OMPS Level-1B (L1B) and Level-2 (L2) products. Through its convenient and enhanced tools/services such as OPeNDAP and L2 Subsetting, GES DISC offers air quality remote sensing user communities facile solutions for complex Earth science data and applications. This presentation will demonstrate TROPOMI and OMPS products including earthview radiance, solar irradiance, and currently available L2 datasets, as well as easy ways to access, visualize and subset data. The implementation of the End User License Agreement (EULA) between NASA GES DISC and all data users accessing data at GES DISC will be emphasized as well.

TROPOMI↗

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj↗

Aerodynamic Sensitivities over Separable Shape Tensors

Here, we present a comprehensive aerodynamic sensitivity analysis of airfoil parameterization informed by separable shape tensors. This parameterization approach uniquely benefits the design process by isolating various well-studied shape characteristics, such as airfoil thickness, and providing a well-regulated low-dimensional parameter domain for aerodynamic designs. Exploring the aerodynamic sensitivities of this novel parameterization can provide valuable insights for more robust designs and future manufacturing efforts. We construct a data-driven parameter space of airfoils using principal geodesic analysis of separable shape tensors informed by a curated database containing almost 20,000 suitable engineering airfoils. Analyzing the shape reconstruction error and the maximum mean discrepancy between joint distributions of aerodynamic quantities, we study the dimensionality of the learned parameter space. This simple numerical experiment demonstrates a dramatic dimension reduction that retains design effectiveness and promotes regularity of the shape representations. Finally, we generate new airfoils and use the HAM2D Reynolds-averaged Navier–Stokes solver to predict lift, drag, and moment coefficients. We compute multiple sensitivity metrics to quantify and assert the consistency of parameter influence on the aerodynamic quantities. We also explore low-dimensional polynomial ridge approximations to motivate physical intuitions and offer explanations of the approximated sensitivities.

17 WIND ENERGY↗

Nasa's Astromaterials Collections: Housed at the NASA Johnson Space Center (JSC) in Houston, TX

The Astromaterials Research and Exploration Science Division at JSC is responsible for the curation of extraterrestrial samples from NASA's past, present and future sample return missions. These samples provide data that help scientists better understand the history and evolution of our Solar System. Our mission is to preserve, protect, and distribute samples for research by the present and future scientific community.

Graff, Paige V.↗

GeoLab 2011: New Instruments and Operations Tested at Desert RATS

GeoLab is a geological laboratory and testbed designed for supporting geoscience activities during NASA's analog demonstrations. Scientists at NASA's Johnson Space Center built GeoLab as part of a technology project to aid the development of science operational concepts for future planetary surface missions [1, 2, 3]. It is integrated into NASA's Habitat Demonstration Unit, a first generation exploration habitat test article. As a prototype workstation, GeoLab provides a high fidelity working space for analog mission crewmembers to perform in-situ characterization of geologic samples and communicate their findings with supporting scientists. GeoLab analog operations can provide valuable data for assessing the operational and scientific considerations of surface-based geologic analyses such as preliminary examination of samples collected by astronaut crews [4, 5]. Our analog tests also feed into sample handling and advanced curation operational concepts and procedures that will, ultimately, help ensure that the most critical samples are collected during future exploration on a planetary surface, and aid decisions about sample prioritization, sample handling and return. Data from GeoLab operations also supports science planning during a mission by providing additional detailed geologic information to supporting scientists, helping them make informed decisions about strategies for subsequent sample collection opportunities.

Evans, Cindy A.↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Apollo Lunar Sample Photograph Digitization Project Update

This is an update of the progress of a 4-year data restoration project effort funded by the LASER program to digitize photographs of the Apollo lunar rock samples and create high resolution digital images and undertaken by the Astromaterials Acquisition and Curation Office at JSC [1]. The project is currently in its last year of funding. We also provide an update on the derived products that make use of the digitized photos including the Lunar Sample Catalog and Photo Database[2], Apollo Sample data files for GoogleMoon[3].

Todd, N. S.↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]↗

Generating Cryogenic Lunar Simulants Within the Planetary Exploration & Astromaterials Research Laboratory

The Planetary Exploration & Astromaterials Research Laboratory (PEARL) aims to provides capabilities for the creation of—and research on—cryogenic lunar regolith simulants (CRS) containing surface volatile analytes. CRS represent regolith samples that might be collected from lunar Permanently Shadowed Regions (PSRs) during future missions. Handling CRS requires procedural and engineering development due to extremely cold (i.e. ≤ - 196°C) working temperatures. However, there is an imperative to understand the physical and chemical alteration of samples during collection, transport, and curation processes throughout the Artemis missions. The chemistries that will be encountered within the PSRs need to be understood based upon prior mission data analysis. The ability to utilize these techniques provides future tests that could be relevant to planetary protection applications regarding future robotic sample return.

Lunar↗

The SunPy Project: An Interoperable Ecosystem for Solar Data Analysis

The SunPy Project is a community of scientists and software developers creating an ecosystem of Python packages for solar physics. The project includes the sunpy core package as well as a set of affiliated packages. The sunpy core package provides general purpose tools to access data from different providers, read image and time series data, and transform between commonly used coordinate systems. Affiliated packages perform more specialized tasks that do not fall within the more general scope of the sunpy core package. In this article, we give a high-level overview of the SunPy Project, how it is broader than the sunpy core package, and how the project curates and fosters the affiliated package system. We demonstrate how components of the SunPy ecosystem, including sunpy and several affiliated packages, work together to enable multi-instrument data analysis workflows. We also describe members of the SunPy Project and how the project interacts with the wider solar physics and scientific Python communities. Finally, we discuss the future direction and priorities of the SunPy Project.

Solar physics↗

GLORIA - A Globally Representative Hyperspectral In Situ Dataset for Optical Sensing of Water Quality

The development of algorithms for remote sensing of water quality (RSWQ) requires a large amount of in situ data to account for the bio-geo-optical diversity of inland and coastal waters. The GLObal Reflectance community dataset for Imaging and optical sensing of Aquatic environments (GLORIA) includes 7,572 curated hyperspectral remote sensing reflectance measurements at 1 nm intervals within the 350 to 900 nm wavelength range. In addition, at least one co-located water quality measurement of chlorophyll a , total suspended solids, absorption by dissolved substances, and Secchi depth, is provided. The data were contributed by researchers affiliated with 59 institutions worldwide and come from 450 different water bodies, making GLORIA the de-facto state of knowledge of in situ coastal and inland aquatic optical diversity. Each measurement is documented with comprehensive methodological details, allowing users to evaluate fitness-for-purpose, and providing a reference for practitioners planning similar measurements. We provide open and free access to this dataset with the goal of enabling scientific and technological advancement towards operational regional and global RSWQ monitoring.

remote sensing of water quality↗

NASA GeneLab Project: Bridging Space Radiation Omics with Ground Studies

Accurate assessment of risk factors for long-term space missions is critical for human space exploration: therefore it is essential to have a detailed understanding of the biological effects on humans living and working in deep space. Ionizing radiation from Galactic Cosmic Rays (GCR) is one of the major risk factors factor that will impact health of astronauts on extended missions outside the protective effects of the Earth's magnetic field. Currently there are gaps in our knowledge of the health risks associated with chronic low dose, low dose rate ionizing radiation, specifically ions associated with high (H) atomic number (Z) and energy (E). The GeneLab project (genelab.nasa.gov) aims to provide a detailed library of Omics datasets associated with biological samples exposed to HZE. The GeneLab Data System (GLDS) currently includes datasets from both spaceflight and ground-based studies, a majority of which involve exposure to ionizing radiation. In addition to detailed information for ground-based studies, we are in the process of adding detailed, curated dosimetry information for spaceflight missions. GeneLab is the first comprehensive Omics database for space related research from which an investigator can generate hypotheses to direct future experiments utilizing both ground and space biological radiation data. In addition to previously acquired data, the GLDS is continually expanding as Omics related data are generated by the space life sciences community. Here we provide a brief summary of space radiation related data available at GeneLab.

Genelab↗

Antarctic meteorite newsletter. Volume 4: Number 1, February 1981: Antarctic meteorite descriptions, 1976, 1977, 1978, 1979

This issue of the Newsletter is essentially a catalog of all antarctic meteorites in the collections of the Johnson Space Center Curation Facility and the Smithsonian except for 288 pebbles now being classed. It includes listings of all previously distributed data sheets plus a number of new ones for 1979. Indexes of samples include meteorite name/number, classification, and weathering category. Separate indexes list type 3 and 4 chondrites, all irons, all achondrites, and all carbonaceous chondrites.

Stone, R.↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

NASA Johnson Space Center's Planetary Sample Analysis and Mission Science (PSAMS) Laboratory: A National Facility for Planetary Research

NASA Johnson Space Center's (JSC's) Astromaterials Research and Exploration Science (ARES) Division, part of the Exploration Integration and Science Directorate, houses a unique combination of laboratories and other assets for conducting cutting edge planetary research. These facilities have been accessed for decades by outside scientists, most at no cost and on an informal basis. ARES has thus provided substantial leverage to many past and ongoing science projects at the national and international level. Here we propose to formalize that support via an ARES/JSC Plane-tary Sample Analysis and Mission Science Laboratory (PSAMS Lab). We maintain three major research capa-bilities: astromaterial sample analysis, planetary process simulation, and robotic-mission analog research. ARES scientists also support planning for eventual human ex-ploration missions, including astronaut geological training. We outline our facility's capabilities and its potential service to the community at large which, taken together with longstanding ARES experience and expertise in curation and in applied mission science, enable multi-disciplinary planetary research possible at no other institution. Comprehensive campaigns incorporating sample data, experimental constraints, and mission science data can be conducted under one roof.

Draper, D. S.↗

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗