Search NASASearch

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Use of Semantic Technology to Create Curated Data Albums

One of the continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available online. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the data sets they need can obtain the specific files using these systems. However, in cases where researchers are interested in studying an event of research interest, they must manually assemble a variety of relevant data sets by searching the different distributed data systems. Consequently, there is a need to design and build specialized search and discover tools in Earth science that can filter through large volumes of distributed online data and information and only aggregate the relevant resources needed to support climatology and case studies. This paper presents a specialized search and discovery tool that automatically creates curated Data Albums. The tool was designed to enable key elements of the search process such as dynamic interaction and sense-making. The tool supports dynamic interaction via different modes of interactivity and visual presentation of information. The compilation of information and data into a Data Album is analogous to a shoebox within the sense-making framework. This tool automates most of the tedious information/data gathering tasks for researchers. Data curation by the tool is achieved via an ontology-based, relevancy ranking algorithm that filters out nonrelevant information and data. The curation enables better search results as compared to the simple keyword searches provided by existing data systems in Earth science.

Ramachandran, Rahul

Use of Semantic Technology to Create Curated Data Albums

One of the continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available online. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the data sets they need can obtain the specific files using these systems. However, in cases where researchers are interested in studying an event of research interest, they must manually assemble a variety of relevant data sets by searching the different distributed data systems. Consequently, there is a need to design and build specialized search and discovery tools in Earth science that can filter through large volumes of distributed online data and information and only aggregate the relevant resources needed to support climatology and case studies. This paper presents a specialized search and discovery tool that automatically creates curated Data Albums. The tool was designed to enable key elements of the search process such as dynamic interaction and sense-making. The tool supports dynamic interaction via different modes of interactivity and visual presentation of information. The compilation of information and data into a Data Album is analogous to a shoebox within the sense-making framework. This tool automates most of the tedious information/data gathering tasks for researchers. Data curation by the tool is achieved via an ontology-based, relevancy ranking algorithm that filters out non-relevant information and data. The curation enables better search results as compared to the simple keyword searches provided by existing data systems in Earth science.

Ramachandran, Rahul

Improving the Acquisition and Management of Sample Curation Data

This paper discusses the current sample documentation processes used during and after a mission, examines the challenges and special considerations needed for designing effective sample curation data systems, and looks at the results of a simulated sample result mission and the lessons learned from this simulation. In addition, it introduces a new data architecture for an integrated sample Curation data system being implemented at the NASA Astromaterials Acquisition and Curation department and discusses how it improves on existing data management systems.

Todd, Nancy S.

Virtual Collections: An Earth Science Data Curation Service

The role of Earth science data centers has traditionally been to maintain central archives that serve openly available Earth observation data. However, in order to ensure data are as useful as possible to a diverse user community, Earth science data centers must move beyond simply serving as an archive to offering innovative data services to user communities. A virtual collection, the end product of a curation activity that searches, selects, and synthesizes diffuse data and information resources around a specific topic or event, is a data curation service that improves the discoverability, accessibility, and usability of Earth science data and also supports the needs of unanticipated users. Virtual collections minimize the amount of the time and effort needed to begin research by maximizing certainty of reward and by providing a trustworthy source of data for unanticipated users. This presentation will define a virtual collection in the context of an Earth science data center and will highlight a virtual collection case study created at the Global Hydrology Resource Center data center.

Data

Collaborative Data Curation to Support the Multi-Mission Algorithm and Analysis Platform (MAAP)

Upcoming space-borne missions will offer unprecedented data about Earth but will also feature exponentially high data volumes. These high data volumes will change the way the scientific community works with data and will also create a unique need for improved data sharing and collaboration. NASA and ESA are working together to address these issues by collaboratively developing the Multi-Mission Algorithm and Analysis Platform (MAAP) to improve the understanding of global aboveground terrestrial carbon dynamics. The MAAP will support ESA’s BIOMASS mission, NASA’s GEDI mission and NASA/ISRO’s NISAR mission. The MAAP will be developed in two phases: a pilot phase and a full production phase. The pilot phase will demonstrate collaboration and basic capabilities. The pilot phase will focus on biomass relevant airborne and field campaign data. Two NASA teams are supporting the development of the MAAP. The MAAP engineering team is responsible for the development, maintenance and operations of the MAAP system while the MAAP data team ensures the ongoing quality of the data, metadata and other information provided in the MAAP. The MAAP data team also supports the ingest and archive of identified data to the MAAP platform. This poster describes the use case development process for the pilot MAAP and the data curated in support of those use cases. Additionally, this presentation will outline the pilot MAAP data ingest process and metadata curation effort along with efforts to ensure interoperability between ESA and NASA data and metadata.

Bugbee, Kaylin

Aggregation Tool to Create Curated Data albums to Support Disaster Recovery and Response

Despite advances in science and technology of prediction and simulation of natural hazards, losses incurred due to natural disasters keep growing every year. Natural disasters cause more economic losses as compared to anthropogenic disasters. Economic losses due to natural hazards are estimated to be around $6-$10 billion dollars annually for the U.S. and this number keeps increasing every year. This increase has been attributed to population growth and migration to more hazard prone locations such as coasts. As this trend continues, in concert with shifts in weather patterns caused by climate change, it is anticipated that losses associated with natural disasters will keep growing substantially. One of challenges disaster response and recovery analysts face is to quickly find, access and utilize a vast variety of relevant geospatial data collected by different federal agencies such as DoD, NASA, NOAA, EPA, USGS etc. Some examples of these data sets include high spatio-temporal resolution multi/hyperspectral satellite imagery, model prediction outputs from weather models, latest radar scans, measurements from an array of sensor networks such as Integrated Ocean Observing System etc. More often analysts may be familiar with limited, but specific datasets and are often unaware of or unfamiliar with a large quantity of other useful resources. Finding airborne or satellite data useful to a natural disaster event often requires a time consuming search through web pages and data archives. Additional information related to damages, deaths, and injuries requires extensive online searches for news reports and official report summaries. An analyst must also sift through vast amounts of potentially useful digital information captured by the general public such as geo-tagged photos, videos and real time damage updates within twitter feeds. Collecting and aggregating these information fragments can provide useful information in assessing damage in real time and help direct recovery efforts. The search process for the analyst could be made much more efficient and productive if a tool could go beyond a typical search engine and provide not just links to web sites but actual links to specific data relevant to the natural disaster, parse unstructured reports for useful information nuggets, as well as gather other related reports, summaries, news stories, and images. This presentation will describe a semantic aggregation tool developed to address similar problem for Earth Science researchers. This tool provides automated curation, and creates "Data Albums" to support case studies. The generated "Data Albums" are compiled collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; information about the event contained in news reports, and images or videos to supplement research analysis. An ontology-based relevancy-ranking algorithm drives the curation of relevant data sets for a given event. This tool is now being used to generate a catalog of Hurricane Case Studies at Global Hydrology Resource Center (GHRC), one of NASA's Distribute Active Archive Centers. Another instance of the Data Albums tool is currently being created in collaboration with NASA/MSFC's SPoRT Center, which conducts research on unique NASA products and capabilities that can be transitioned to the operational community to solve forecast problems. This new instance focuses on severe weather to support SPoRT researchers in their model evaluation studies

Ramachandran, Rahul

Bringing Research to New Heights: How CASEI Integrates Data Curation, Discovery, and Education in Earth and Atmospheric Science

A challenging aspect of any project is finding all the relevant data and information needed to address the research objective. Searching for data and its contextual metadata can become overwhelming for both undergraduate and graduate students, potentially hindering their work and affecting the scientific discoveries that could be made in the long run. To ease this, the NASA Airborne Data Management Group (ADMG), part of the Interagency Implementation and Advanced Concepts Team (IMPACT), has developed the new Catalog of Archived Suborbital Earth science Investigations (CASEI). CASEI includes a web portal that users, be they professionals or students, can use to search, browse, discover, and locate relevant observations associated with NASA’s airborne and field campaigns. Users are able to query data in a variety of ways (via keywords, locations, timeframe, etc) from one online portal, minimizing the amount of time needed to search. CASEI also allows access to key contextual metadata and data from a wide array of Earth and Atmospheric Science topics such as aerosols and boundary layer processes, as well as ice and glacial properties or processes. Users are able to access the data via DOI links to data set landing pages. This presentation will demonstrate how CASEI can be used for classwork and student research. Teachers can provide CASEI to their students as a tool for their studies, or use it to find data themselves while constructing their curriculums. Additionally, users can leverage CASEI to learn about NASA’s Earth and Atmospheric Science research efforts and to find data relevant for assignments or other research projects. The metadata in CASEI has been carefully curated, and highlights important information about the campaigns and their data. Students can explore and learn about the scientific objectives of the campaigns, as well as descriptions of the campaign’s best research days. Having access to contextual metadata in an easy to understand way can help plant the seeds of new ideas in students at any point in their academic journey. From class projects to theses/dissertations and other research, CASEI is a valuable emerging tool for data discovery, giving access to all users and guiding researchers to NASA’s unique airborne data to answer the burning Earth Science questions of our time.

education

Kaona: Deep Searching and Curating Data from Aviation Safety Reporting Systems

Context: Several works in the literature have examined how safety narrative databases can be leveraged to share lessons learned. However, less attention has been given to augmenting existing processes for mining these safety reporting system databases. Aim: In this work, we introduce Kaona: An interface that weaves machine learning in existing aviation safety database mining activities. Method: We provide a use case of search, curation and newsletter writing to showcase how Kaona features build on existing processes and on its own to enhance information retrieval, curation and synthesis of narratives. Results: We created two instances of Kaona internally for evaluation, one using publicly available NASA’s ASRS narratives and another using publicly available C3RS narratives. Data ranged from 1998 to 2024. Conclusion: Our tool provides a new way to explore safety narratives, serving to re-imagine how text databases can benefit of novel information retrieval mechanisms in the era of large language models.

ASRS

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena

Data Mining for Anomaly Detection

The Vehicle Integrated Prognostics Reasoner (VIPR) program describes methods for enhanced diagnostics as well as a prognostic extension to current state of art Aircraft Diagnostic and Maintenance System (ADMS). VIPR introduced a new anomaly detection function for discovering previously undetected and undocumented situations, where there are clear deviations from nominal behavior. Once a baseline (nominal model of operations) is established, the detection and analysis is split between on-aircraft outlier generation and off-aircraft expert analysis to characterize and classify events that may not have been anticipated by individual system providers. Offline expert analysis is supported by data curation and data mining algorithms that can be applied in the contexts of supervised learning methods and unsupervised learning. In this report, we discuss efficient methods to implement the Kolmogorov complexity measure using compression algorithms, and run a systematic empirical analysis to determine the best compression measure. Our experiments established that the combination of the DZIP compression algorithm and CiDM distance measure provides the best results for capturing relevant properties of time series data encountered in aircraft operations. This combination was used as the basis for developing an unsupervised learning algorithm to define "nominal" flight segments using historical flight segments.

Biswas, Gautam

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert

GHRSST-14 DAS-TAG Report

The DAS-TAG provides the informatics and data management expertise in emerging information technologies for the GHRSST community. It provides expertise in data and metadata formats and standards, fosters improvements for GHRSST data curation, experiments with new data processing paradigms, and evaluates services and tools for data usage. It provides a forum for producer and distributor data management issues and coordination.

data processing

Curating Virtual Data Collections

NASAs Earth Observing System Data and Information System (EOSDIS) contains a rich set of datasets and related services throughout its many elements. As a result, locating all the EOSDIS data and related resources relevant to particular science theme can be daunting. This is largely because EOSDIS data's organizing principle is affected more by the way they are produced than around the expected end use. Virtual collections oriented around science themes can overcome this by presenting collections of data and related resources that are organized around the user's interest, not around the way the data were produced. Virtual collections consist of annotated web addresses (URLs) that point to data and related resource addresses, thus avoiding the need to copy all of the relevant data to a single place. These URL addresses can be consumed by a variety of clients, ranging from basic URL downloaders (wget, curl) and web browsers to sophisticated data analysis programs such as the Integrated Data Viewer.

data integration

Evolution of NASA's Earth Science Digital Object Identifier Registration System

NASA's Earth Science Data and Information System (ESDIS) Project has implemented a fully automated system for assigning Digital Object Identifiers (DOIs) to Earth Science data products being managed by its network of 12 distributed active archive centers (DAACs). A key factor in the successful evolution of the DOI registration system over last 7 years has been the incorporation of community input from three focus groups under the NASA's Earth Science Data System Working Group (ESDSWG). These groups were largely composed of DOI submitters and data curators from the 12 data centers serving the user communities of various science disciplines. The suggestions from these groups were formulated into recommendations for ESDIS consideration and implementation. The ESDIS DOI registration system has evolved to be fully functional with over 5,000 publicly accessible DOIs and over 200 DOIs being held in reserve status until the information required for registration is obtained. The goal is to assign DOIs to the entire 8000+ data collections under ESDIS management via its network of discipline-oriented data centers. DOIs make it easier for researchers to discover and use earth science data and they enable users to provide valid citations for the data they use in research. Also for the researcher wishing to reproduce the results presented in science publications, the DOI can be used to locate the exact data or data products being cited.

DOI; digital object identifier; mappin