Search NASASearch

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Increasing Data Discovery and Re-Use: The Space Life Sciences Ontology

Two of the most important goals of the adoption of the FAIR principles are increasing the ability of agents to find and re-use research data. Achieving these goals for space life sciences research is even more pressing, given the relatively expensive and scarce nature of these data. We have reported in the past on the progress made by exemplar life sciences data systems towards implementing FAIR, showing gaps particularly in the “interoperability area” of the principles; the lack of common conceptual models for space life science research is one reason for this gap. There were few available resources that define, annotate, categorize or otherwise relate various kinds of metadata describing the acquisition, nature, and intent of investigational space life sciences data. To address this gap, NASA is working with the Open Biological and Biomedical Ontology Foundry (https://obofoundry.org/) to develop the Space Life Science Ontology (SLSO) that is intended to support archival and other kinds of systems that operate using these data. The scope of the ontology includes concepts regarding those aspects of investigation design and execution specific or unique to space environments, such as types of specialized equipment, operating organizations, and documentation. The ontology is continually being developed and published to the life science community (https://github.com/nasa/LSDAO/); at the time of this publication, the SLSO newly and uniquely defines 30 types (classes), 90 properties, and 14 relations specific to space life sciences metadata. In addition, the SLSO reuses (imports) some 2,360 types (classes), 49 properties, and 393 relations from other ontologies that are relevant to these kinds of metadata. In addition to its role as a common conceptualization for space biomedical research activities, the SLSO can also be used to provide automated support for traditionally difficult and expensive activities such as data curation and cross-system data integration and analysis.

fair

Open Science for Plants in Space: Improvements in NASA's Open Science Data Repository

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, elevated CO2, and many other abiotic stressors. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. Current OSDR standards include the ISA (Investigation-Study-Assay) experiment model, assay metadata configurations, and standardized terminology and ontologies. In 2024 OSDR will include a new suite of features for improved FAIR compliance including downloadable plant metadata templates, data submission tools and overall improved AI-readiness of plant datasets. AI/ML methods can be helpful tools to overcome the inherent challenges of space biology research (small sample size, sparse and heterogeneous data etc.). However these methods are built on an assumption of normalized and well-curated data. OSDR’s new curation tools will improve users ability to leverage ML and AI methods to model space biology data and better understand the complex effects of spaceflight on living systems across hierarchical biological levels. We look forward to sharing our advances with the spaceflight community.

FAIR

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott

Database Design Strategies for Coordinated Simulation and Testing in Additive Manufacturing

The qualification and certification (Q&C) process presents a significant challenge for widespread adoption of additive manufacturing (AM) materials and processes for aerospace applications. A relational database framework will be presented as a tool for data curation of coordinated experimental and computational materials modeling research activities. A comparison of relational and hierarchical data structures in this domain will be emphasized through the evolution of a database design strategy. This framework’s mission is to support the advancement of computational materials-informed Q&C by providing the necessary data infrastructure to trace reliability and reproducibility measures through unified AM materials simulation and experimental testing. FAIR (findable, accessible, interoperable, and reusable) data will be highlighted as a necessary precursor for automation of specific actions, which ultimately reduces the time and expense burden for Q&C. The discussion will be mostly limited to back-end design elements, though a few front-end user experience examples will also be shared.

Qualification

Making Heliophysics Research More Open and Accessible at the Community Coordinated Modeling Center (CCMC)

The Space Weather and Heliophysics modeling community seeks to improve our understanding of space weather events and their impact on human activities. The Community Coordinated Modeling Center’s (CCMC, https://ccmc.gsfc.nasa.gov) mission is to support the community by providing a convenient collaborative platform that brings together space weather models, model simulation data, curated datasets of solar events, and associated value-added services. Using these services, researchers and other end-users may exercise, evaluate, and intercompare contributed models, triage designated R2O models, as well as collaborate on a continuously updated archive of model run results. This presentation reports on CCMC’s ongoing efforts in making Heliophysics models and data more accessible, open, and reproducible. We will also explore interoperability within the ecosystem of CCMC services and how this ecosystem interoperates with external partner services and data streams.

space weather

Populating a Graph Database to Run a Usage-Based Discovery Tool

Most dataset discovery tools for Earth Observation data rely on descriptions and other metadata of the datasets, using keyword searches or attribute filtering to determine relevance. However, these descriptions often do not include the potential uses of the data. Thus, a user working on floods will rarely see few if any rainfall datasets show up in such a search. The Usage Based Discovery tool, on the other hand, offers usage instances to the user, either research articles or applications, along with the datasets that those usage instances used. This allows a user, particularly one new to the world of Earth Observation data, to investigate which datasets are used in similar cases. The information that powers Usage-Based Discovery is a graph database of relationships of usage to dataset and usage to topic, allowing the user to narrow their search for similar cases. In order to scale out to a graph database rich enough to provide a satisfactory user experience, we combine manual and automated processes to populate the graph. The initial content of the graph has been seeded primarily via human-aided data curation methods, using sites like Google Scholar. To scale up this effort, we’ve employed crowdsourcing. It is easy for anyone to contribute to our graph using their Open Researcher and Contributor Identifier for authorization. We’re now experimenting with Machine Learning and Natural Language Processing to help automate population of the graph, starting with the classification of research articles by topic. Finding adequate training data in the absence of a comprehensive and open research article API continues to be a significant challenge.

Vincent Inverso

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors

Lowering Barriers to Science and Space Weather Research at the Community Coordinated Modeling Center (CCMC)

The Space Weather and Heliophysics research and modeling community has been pushing the limits of our ability to understand and predict space weather events. The Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) supports the community by providing a convenient collaborative platform hosting space weather models, model simulation data, curated datasets of solar events, and associated value-added services. Using these services, researchers and other end-users may exercise, evaluate, and intercompare contributed models, triage designated R2O models, as well as collaborate on a continuously updated archive of model run results. We will focus on CCMC’s ongoing commitment to the principles and guidelines of the Open Science initiative. Particularly, we will discuss our work towards making our services more transparent and our library of model simulations more accessible, open, and reproducible. We will introduce our recent tools for data discovery and correlative analysis designed to further increase the value of the user-generated data and metadata. We will also present our recent work on making heliophysical models more accessible and open to the community, particularly through simplified user experience and expert domain support. We will report on our progress in establishing an inter-center infrastructure with the ESA Virtual Space Weather Modelling Centre (VSWMC), designed to cross organizational boundaries and provide streamlined access to a joint palette of the models.

space weather

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical

Lunar Glovebox Balance with Wireless Technology

The most important equipment required for processing lunar samples is a high-quality mass balance for maintaining accurate weight inventory, security, and scientific study. After careful review, a Curation Office memo by Michael Duke in 1978 chose the Mettler PL200 to be used for sample weight measurements inside the gloveboxes (Fig. 3). These commercial off-the-shelf (COTS) balances did not meet the strict accepted material requirements in the Lunar lab. As a result, each balance housing, weighing pan, and wiring was custom retrofitted to meet Lunar Operating Procedure (LOP) 54 requirements [for material construction restrictions]. The original design drawings for the custom housings, readout support stands, and wiring were done by the JSC engineering directorate. The 1977- 1978 schematics, drawings, and files are now housed in the curation Data Center. Per the design specifications, the housing was fabricated from aluminum grade 6061 T6, seamless welds, and anodized per MIL-A-8625 type I, class I. The balance feet were TFE Teflon and any required joints were sealed with Viton A gaskets. The readout display and support stands outside the glovebox were fabricated from 300 series stainless steel with #4 finish and mounted to the glovebox with welded bolts. Wire harnesses that linked the balance with the outside display and power were encapsulated with TFE Teflon and transported through custom Deutsch wire bulk head pass-through systems from inside to outside the glovebox. These Deutsch connectors were custom fabricated with 316L stainless steel bodies, Viton A O-rings, aluminum 6061 with electroless nickel plating, Teflon (replacing the silicone), and gold crimp connectors (no soldering). Many of the Deutsch connectors may have been used in the Apollo program high vacuum complex in building 37 and date to about 1968 to 1970.

Zeigler, Ryan A.

Extravehicular Activity Mission System Software (EMSS) - Enabling Human Planetary Exploration Data Within The Broader Planetary Data Ecosystem

The planetary science community is once again on the verge of generating, capturing and analyzing human planetary exploration data, this time via the Artemis program. Artemis missions will involve robotic missions in addition to human extravehicular activity (EVA) where crew will be generating scientific data [1]. Present-day robotic mission data expectations for data archiving involves ingesting data into the Planetary Data System (PDS), but how might PDS be leveraged/adapted/ready (or not) for human spaceflight mission data, particularly EVA data that includes non-scientific data that provides important context to the scientific data gathered on the lunar surface? This question has broader implications than what this abstract can answer, but we wanted to pose the question to 1) get conversations started and 2) highlight how operations software data handling could play a role in overall data curation.

M J Miller

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship. 1. Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. 2. National Academies of Sciences, E. and Medicine, Open Science by Design: Realizing a Vision for 21st Century Research. 2018, Washington, DC: The National Academies Press. 232. 3. Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5.

Life Sciences data

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.

Life Sciences data

NLSP: NASA Life Sciences Portal

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles. Some of these improvements will at the same time support the twin pillars of Open Science: transparency of methods and reproducibility of results. This video is a high level overview of the NLSP for existing and new users.

Life Sciences data

Technical Tension Between Achieving Particulate and Molecular Organic Environmental Cleanliness: Data from Astromaterial Curation Laboratories

NASA Johnson Space Center operates clean curation facilities for Apollo lunar, Antarctic meteorite, stratospheric cosmic dust, Stardust comet and Genesis solar wind samples. Each of these collections is curated separately due unique requirements. The purpose of this abstract is to highlight the technical tensions between providing particulate cleanliness and molecular cleanliness, illustrated using data from curation laboratories. Strict control of three components are required for curating samples cleanly: a clean environment; clean containers and tools that touch samples; and use of non-shedding materials of cleanable chemistry and smooth surface finish. This abstract focuses on environmental cleanliness and the technical tension between achieving particulate and molecular cleanliness. An environment in which a sample is manipulated or stored can be a room, an enclosed glovebox (or robotic isolation chamber) or an individual sample container.

Allton, J. H.

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado