Search NASA⌕ Search

SEARCH · Search NASA

Results for “dataset curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Curating AI-Ready Datasets for Equity and Environmental Justice: A Data-Centric AI Case Study

An equitable and environmentally just community is essentialin order to avoid disproportionate burden borne by vulnerablecommunities. This need becomes pressing in the aftermathof an extreme event such as disaster or hazard when it is diffi-cult for the governing bodies to implement resource allocationas per the need. Artificial Intelligence (AI) algorithms canhelp surface Equity and Environmental Justice (EEJ) issueswhen trained on EEJ datasets. However, curating AI-readyEEJ training datasets is challenging due to differences in fac-tors such as heterogeneity, resolution, modality, and level ofexpertise in labeling. Additionally, EEJ issues involve sensi-tive information where uncertainties and errors could degradethe performance of AI algorithms. For eg. Error in seasonalcrop yield information can highly affect the prediction of an-nual crop yield. To address these challenges, Data-centricAI (DCAI) methods are employed, which enhance AI algo-rithm performance even with limited training samples. DCAIprioritizes data quality, thereby reducing the adverse effectsof uncertainties and errors during the model training process.This research proposes a novel dataset and benchmark for an-alyzing the effect of the Maui Wildfire of 2023 for Equityand Environmental Justice (EEJ) issues. The proposed datasetaligns with the concepts of DCAI such as annotation quality,data preprocessing, privacy, feature engineering, governanceand provenance. We firmly believe that the proposed datasetwould lay a foundation to implement robust and reliable mod-ern AI algorithms for addressing EEJ issues.

Paridhi Parajuli↗

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors↗

GES DISC Datalist Improves Earth Science Data Discoverability

At American Geophysical Union(AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a novel way to access data: Datalist. Currently, datalist is a collection of predefined data variables from one or more archived datasets, curated by our subject matter expert (SME). Our science support team has curated a predefined Hurricane Datalist and received very positive feedback from the user community. Datalist uses the same architecture our new website uses and have the same look and feel as other datasets on our web site. and also provides a one-stop shopping for data, metadata, citation, documentation, visualization and other available services. Since the last AGU Meeting, we have further developed a few new datalists corresponding to the Big Earth Data Initiative (BEDI) Societal Benefit Areas and A-Train data. We now have four datalists: Hurricane, Wind Energy, Greenhouse Gas and A-Train. We have also started working with our User Working Group members to create their favorite datalists and working with other DAAC to explore the possibility to include their products in our datalists that may also lead to a future of potential federated (cross-DAAC) datalists. Since our datalist prototype effort was a success, we are planning to make datalist operational. It's extremely important to have a common metadata model to support datalist, this will also be the foundation of federated datalist. We mapped our datalist metadata model to the unpublished UMM(Universal Metadata Model)-Var (Variable) (June version) and found that the UMM-var together with UMM-C (Collection) and possible UMM-S (Service) will meet our basic requirements. For example: Dataset shortname, and version are already specified in UMM-C, variable name, long name, units, dimensions are all specified in UMM-Var. UMM-Var also facilitates Science Keywords to allow tagging at variable level and Characteristics for optional variable characteristics. Measurements is useful for grouping of the variables and Set is promising to define datalist. And finally, the UMM-Service model to specify the available services for the variable will be very beneficial. In summary, UMM-Var, UMM-C and UMM-S are the basis of federated datalist and the development and deployment of datalist will contribute to the evolution of the UMM.

datalist↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Making Heliophysics Research More Open and Accessible at the Community Coordinated Modeling Center (CCMC)

The Space Weather and Heliophysics modeling community seeks to improve our understanding of space weather events and their impact on human activities. The Community Coordinated Modeling Center’s (CCMC, https://ccmc.gsfc.nasa.gov) mission is to support the community by providing a convenient collaborative platform that brings together space weather models, model simulation data, curated datasets of solar events, and associated value-added services. Using these services, researchers and other end-users may exercise, evaluate, and intercompare contributed models, triage designated R2O models, as well as collaborate on a continuously updated archive of model run results. This presentation reports on CCMC’s ongoing efforts in making Heliophysics models and data more accessible, open, and reproducible. We will also explore interoperability within the ecosystem of CCMC services and how this ecosystem interoperates with external partner services and data streams.

space weather↗

Lowering Barriers to Science and Space Weather Research at the Community Coordinated Modeling Center (CCMC)

The Space Weather and Heliophysics research and modeling community has been pushing the limits of our ability to understand and predict space weather events. The Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) supports the community by providing a convenient collaborative platform hosting space weather models, model simulation data, curated datasets of solar events, and associated value-added services. Using these services, researchers and other end-users may exercise, evaluate, and intercompare contributed models, triage designated R2O models, as well as collaborate on a continuously updated archive of model run results. We will focus on CCMC’s ongoing commitment to the principles and guidelines of the Open Science initiative. Particularly, we will discuss our work towards making our services more transparent and our library of model simulations more accessible, open, and reproducible. We will introduce our recent tools for data discovery and correlative analysis designed to further increase the value of the user-generated data and metadata. We will also present our recent work on making heliophysical models more accessible and open to the community, particularly through simplified user experience and expert domain support. We will report on our progress in establishing an inter-center infrastructure with the ESA Virtual Space Weather Modelling Centre (VSWMC), designed to cross organizational boundaries and provide streamlined access to a joint palette of the models.

space weather↗

Curating a Standardized Dataset for Statistical Biosignature Classification

In recent years, machine learning has been explored as a toolkit for planetary science and operations [Helbert, Azari]. Machine learning has been used to improve our understanding of possible biosignatures and mineral signatures to improve science return on future missions [Warren-Rhodes, Cleaves].

Biosignatures↗

Changing Climate, Changing Data: Exposing Climate Data to New Users Through GeoPlatform.gov’s Resilience Community

Over 700 climate related datasets were curated by subject matter experts into 9 thematic areas as a part of the Climate Data Initiative (CDI). NASA was tasked with maintaining the collection’s data inventory and supporting web pages at data.gov/climate. Today, the Data Curation for Discovery (DCD) team at MSFC continues to support the CDI collection. In order to expose the collection to a new and growing user community, the DCD team has partnered with GeoPlatform.gov to develop the Resilience community. The Resilience community serves as an interactive, topically-focused web portal that further promotes and shares CDI web content, datasets, services, maps, and other tools relevant to global resilience and change. This poster focuses on the team’s efforts to leverage GeoPlatform’s semantic applications to link CDI objects within the platform to improve discoverability. This poster also provides insights as to how this effort may serve as an example for building and expanding future Geoplatform.gov communities..

Sisco, Adam↗

Citizen Science Approach for Searching and Curating Literature of the Effects of Spaceflight on Cardiovascular Outcomes in Rodents and Humans

The spaceflight environment causes significant changes to the structure and function of the cardiovascular system, including fluid redistribution, alterations in blood pressure, and changes in cardiac output. The goal of this project is to quantitatively summarize the data on the effects of actual or simulated microgravity and radiation exposure resulting from spaceflight on the cardiovascular system. As the first step, a group of investigators approached through a collaboration of the Ames Life Science Data Archive (ALSDA) Analysis Working Group developed a list of relevant cardiovascular search terms. Based on these, medical librarians generated and executed the search strategy in Medline, CINAHL, Embase and NASA repositories. In parallel, we recruited students and young professionals from various space industry-affiliated organizations, resulting in ~100 individuals joining. With this program we aimed to reach students and young people underrepresented in STEM, including first-generation, female, minorities, disadvantaged backgrounds, fostered individuals, etc. These individuals completed a virtual training course on the nature and methodologies of the project. Following this, the participants were structured into teams with more senior/experienced individuals designated as team leaders. Currently, the teams are screening approximately 15,000 studies using the systematic review tool, Covidence. Teams will be extracting and curating data for meta-analysis of the cardiovascular spaceflight literature, but also extracting, submitting, and curating appropriate datasets into the new ALSDA submission portal and repository. This effort will result in collaborative publications based upon the literature meta-analyses, and a number of publicly accessible datasets for reuse, modeling, machine learning, and knowledge graph-type approaches. This approach reduces the length of time to complete title/abstract screening time from 1-2 years needed for this volume of studies, to 3-4 months, while also providing a unique, open-access educational experience to space research and training in knowledge synthesis tools to interested individuals.

space biology↗

The Historical Greenland Climate Network (GC-Net) Curated and Augmented Level-1 Dataset

The Greenland Climate Network (GC-Net) consists of 31 automatic weather stations (AWSs) at 30 sites across the Greenland Ice Sheet. The first site was initiated in 1990, and the project has operated almost continuously since 1995 under the leadership of the late Konrad Steffen. The GC-Net AWS measured air temperature, relative humidity, wind speed, atmospheric pressure, downward and reflected shortwave irradiance, net radiation, and ice and firn temperatures. The majority of the GC-Net sites were located in the ice sheet accumulation area (17 AWSs), while 11 AWSs were located in the ablation area, and two sites (three AWSs) were located close to the equilibrium line altitude. Additionally, three AWSs of similar design to the GC-Net AWS were installed by Konrad Steffen's team on the Larsen C ice shelf, Antarctica. After more than 3 decades of operation, the GC-Net AWSs are being decommissioned and replaced by new AWSs operated by the Geological Survey of Denmark and Greenland (GEUS). Therefore, making a reassessment of the historical GC-Net AWS data is necessary. We present a full reprocessing of the historical GC-Net AWS dataset with increased attention to the filtering of erroneous measurements, data correction and derivation of additional variables: continuous surface height, instrument heights, surface albedo, turbulent heat fluxes, and 10 m ice and firn temperatures. This new augmented GC-Net level-1 (L1) AWS dataset is now available at https://doi.org/10.22008/FK2/VVXGUT (Steffen et al., 2023) and will continue to be refined. The processing scripts, latest data and a data user forum are available at https://github.com/GEUS-Glaciology-and-Climate/GC-Net-level-1-data-processing (last access: 30 November 2023). In addition to the AWS data, a comprehensive compilation of valuable metadata is provided: maintenance reports, yearly pictures of the stations and the station positions through time. This unique dataset provides more than 320 station years of high-quality atmospheric data and is available following FAIR (findable, accessible, interoperable, reusable) data and code practices.

Greenland Climate Network↗

Ongoing Work: A Prototype Dataset for Low-flying Autonomous Medical UAS Operations

This paper presents ongoing work to create a dataset for low-flying autonomous medical UAS operations, focused on human stance recognition. This is an exploration of the viability of airborne classification for the Drone as a First Responder (DFR) concept in which a UAS arrives at the scene of an incident before emergency response personnel can get there and provides some level of situational awareness for the personnel arriving to the scene. Future incarnations could also see the UAS administer some level of care to injured parties at the scene. The data set, focused on detecting human stance, being developed here is the result of 30 test flights at NASA Langley Research Center in early 2024. In addition to flights where the participant (an anthropomorphic testing device or human) is alone in the viewing area holding a particular stance, two emergency scenes have been fabricated and collected through video - ``bike crash'' and ``difficult camping''. These test flights include four human participants. The contribution of this work upon completion will be a publicly available data set for the development of classification engines focused on human stance, and in the future, even triage.

Uncrewed Aerial Systems↗

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian↗

Analyzing Natural Language Context in Human-Machine Teaming using Supervised Machine Learning

Building a foundation for trustworthiness and trust verification in multi-asset teaming is the research challenge of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR). The Design Reference Mission (DRM) for ATTRACTOR is a search and rescue mission objective governed by a multi-member team consisting of human and machine operators. A crucial component to the effort is the communication between humans and autonomous agents throughout both planning and execution stages of the mission. Intuitive communication methods and modalities are posited as critical enablers for certifying trust and trustworthiness. This paper reports on the data collection and analysis conducted in support of the Human Informed Natural-language GANs Evaluation (HINGE)project to attain explainable and trusted communication between human-machine assets. Two identically curated image description datasets were acquired for HINGE, both consisting of two unique input modalities (typed vs. verbal) and retrieved in two distinct contexts (general vs. specific). The gathered datasets were assessed and compared using Parts-of-Speech (POS)features, sentence similarity metrics, and linguistic analysis. Then, the datasets were modeled and tested separately and in combination with one another using machine learning algorithms. The comparison and testing results reveal a superior dataset, by which a preferred context and input is understood, for generating image representations of missing persons using a Generative Adversarial Network (GAN).

Bryan A Barrows↗

Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) Datasets Released by NASA GES DISC and Their Applications for Air Quality

Nitrogen dioxide (NO2), a pervasive air pollutant, comes from vehicles, power plants, industrial emissions, and off-road sources such as construction or lawn and gardening equipment. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) curates many remote sensing datasets with NO2 retrievals, which have been utilized for air quality research and applications. The remotely-sensed datasets include those generated by the Ozone Monitoring Instrument (OMI) on the Aura satellite, the TROPOspheric Monitoring Instrument (TROPOMI) onboard the Copernicus Sentinel-5 Precursor (S5P), and the Ozone Mapping and Profiling Suite (OMPS) Nadir-Mapper (NM) instrument on the Suomi National Polar-orbiting Partnership (S- NPP). In collaboration with the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) project, the GES DISC recently released MINDS datasets. The NASA MEaSUREs MINDS project aims to develop long-term NO2 global data records by adapting a consistent retrieval algorithm to multiple instrument measurements. Long-term data records will be achieved by applying consistent retrieval approaches to multiple satellite instruments, including OMI (2004 - ); the Global Ozone Monitoring Experiment (GOME, 1995-2011) onboard the second European Remote Sensing satellite (ERS-2); the Scanning Imaging Spectrometer for Atmospheric Cartography (SCIAMACHY, 2002-2012) onboard the ENVIronmental SATellite (ENVISAT); GOME-2 on the Meteorological Operational satellites (MetOp-A and MetOp-B, 2006 - ); and TROPOMI onboard the Copernicus S5P (2017 - ). The long-term record (1995 to present) of MINDS datasets makes them very useful for air quality trend studies. Some MINDS datasets with high spatial resolution of only a few kilometers can be used for air quality research and applications at regional scales. In this presentation, we will introduce all of the MINDS products and services, and demonstrate use cases of MINDS data for studying air quality. We will also present a few other NO2 datasets acquired from NASA’s Health and Air Quality Applied Sciences Team (HAQAST), to be archived and distributed by the GES DISC, and highlight some of their applications for air quality and health.

Feng Ding↗

Enabling API Access to the Space Weather Services at the Community Coordinated Modeling Center

Over the span of 20 years, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) has been leading a number of community-driven services and applications that provide a free and open access to the cutting-edge space weather and Heliophysics models through a simple web-based interface. CCMC also oversees an open archive of user model simulations and related metadata, maintains space weather-related data streams, curates validation and event datasets, and more. To maximize utility of its complementary services and data holdings, CCMC has been gradually building up an ad-hoc set of interfaces and specifications that facilitate coupling and interconnection within the organization, while also simplifying management and monitoring of the data. As the models continue to grow in maturity and complexity, CCMC is looking to reduce the complexity for its end users by making public some of the internal APIs as well as implementing dedicated interfaces as required by the community. In this presentation, we will overview the current run services at CCMC and will describe our current and near-future efforts in providing interfaces to these services.

space weather↗