Search NASASearch

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Configuration and Implementation

The main objectives of a Planetary Data System (PDS) are to curate planetary data - that is to store and maintain complete planetary data sets and the necessary supporting documentation to make them useful - and to facilitate scientific study of these data through improved organization and access. A hardware system, no matter how sophisticated and supportive, cannot replace the live colleague interaction needed for good science.System users with only a general knowledge of a data file or data set will need guidance in determining whether that data is most appropriate to support a theory. The PDS is meant to allow easy access to data, to provide useful presentations of that data, and to facilitate data analysis; however it is not designed to supercede the critical examination and interactions that data analysis requires. The Planetary Data System will be a powerful research tool which allows the entire planetary science community access to a complete data and will facilitate the archiving of pass, current, and future mission data sets.

Source record

Managing Multi-Instrument Data Streams in Secure Environments

The capture and curation of all primary instrument data is a potentially valuable source of added insight into experiments or diagnostics in laboratory experiments. The data can, when properly curated, enable analysis beyond the current practice that uses just a subset of the as-measured data. Complete curated data can also be input for machine learning and other data exploration tools. Conveniently storing and accessing instrument data requires that the instruments are connected to databases and users through a networking infrastructure. This infrastructure needs to accommodate a wide array of instruments which can range from single laboratory mounted probes for environment monitoring to computers managing multiple instruments. These resources may also include mobile devices on which researchers record instrument and experiment state related notes. These varied data sources bring with them the challenges of different communications capabilities and protocols as well as the primary data typically being produced in proprietary formats. These challenges are further compounded when the instruments need to operate in secure environments such as required in national laboratories. We will discuss the SmartLab, an ongoing effort to set up a system for instrument and simulation data curation at NASA Langley Research Center. We will outline the challenges faced in managing the data sources required for ongoing research activities and the solutions that are being considered and implemented to address those challenges.

instrument data management

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

Approaches used in Earth science research such as case study analysis and climatology studies involve discovering and gathering diverse data sets and information to support the research goals. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. In cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. This paper presents a specialized search, aggregation and curation tool for Earth science to address these challenges. The search rool automatically creates curated 'Data Albums', aggregated collections of information related to a specific event, containing links to relevant data files [granules] from different instruments, tools and services for visualization and analysis, and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non relevant information and data.

Ramachandran, Rahul

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

One of the largest continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available. Approaches used in Earth science research such as case study analysis and climatology studies involve gathering discovering and gathering diverse data sets and information to support the research goals. Research based on case studies involves a detailed description of specific weather events using data from different sources, to characterize physical processes in play for a specific event. Climatology-based research tends to focus on the representativeness of a given event, by studying the characteristics and distribution of a large number of events. This allows researchers to generalize characteristics such as spatio-temporal distribution, intensity, annual cycle, duration, etc. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the datasets of interest can obtain the specific files they need using these systems. However, in cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. In these cases, a search process needs to be organized around the event rather than observing instruments. In addition, the existing data systems assume users have sufficient knowledge regarding the domain vocabulary to be able to effectively utilize their catalogs. These systems do not support new or interdisciplinary researchers who may be unfamiliar with the domain terminology. This paper presents a specialized search, aggregation and curation tool for Earth science to address these existing challenges. The search tool automatically creates curated "Data Albums", aggregated collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non-relevant information and data.

Ramachandran, Rahul

Visualizing Geospatial Data through ESRI Story Maps for Earth Science Education: Lessons Learned from My NASA Data

For 20 years My NASA Data (MND) has curated NASA Earth science data and provided the data to educators in engaging learner-centered resources. MND has recently featured story maps as an innovative way to engage students in NASA Earth data. A story map is a cloud-based lesson that engages the learner in interactive geospatial maps using NASA data, and other multimedia content, text, and tasks that can be seamlessly incorporated in classroom instruction. This immersive technology eliminates the need for the user to move among tricky interfaces to access and visualize Earth science data, and no special software is required to be downloaded. Each story map integrates data from different NASA satellite missions, retrieved from Distributed Active Archive Centers (DAACs). Story maps also employ data analysis tools, such as time series options and swipe tools that allow learners to view and analyze relationships between scientific variables. MND has produced 25 story map lesson plans on the topics of air quality, the urban heat island effect, Earth’s energy budget, phytoplankton distribution, hurricane formation, solar eclipses, ocean circulation patterns, sea ice extent, and volcanic eruptions. Nine of them are extended story maps and written in the 5E format, which is internationally recognized as best practice based on how children learn science. Each story map resource is developed by the MND team featuring a GIS programming specialist, a lead scientist, and educational specialist/s to ensure the context, content, and methods are scientifically and educationally sound. The MND story maps are written for middle and high school science teachers and students as they connect with the Earth Systems Science phenomena featured in the Next Generation Science Standards. Each story map includes supporting resources for smooth integration in the classroom. During Fiscal Year 2023, The My NASA Data website received over 100,000 story map engagements during. These metrics highlight the interest in story maps as an Earth Science educational resource.

Desiray Wilson

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism, behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology

Enabling API Access to the Space Weather Services at the Community Coordinated Modeling Center

Over the span of 20 years, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) has been leading a number of community-driven services and applications that provide a free and open access to the cutting-edge space weather and Heliophysics models through a simple web-based interface. CCMC also oversees an open archive of user model simulations and related metadata, maintains space weather-related data streams, curates validation and event datasets, and more. To maximize utility of its complementary services and data holdings, CCMC has been gradually building up an ad-hoc set of interfaces and specifications that facilitate coupling and interconnection within the organization, while also simplifying management and monitoring of the data. As the models continue to grow in maturity and complexity, CCMC is looking to reduce the complexity for its end users by making public some of the internal APIs as well as implementing dedicated interfaces as required by the community. In this presentation, we will overview the current run services at CCMC and will describe our current and near-future efforts in providing interfaces to these services.

space weather

Geocuration Lessons Learned from the Climate Data Initiative Project

Curation is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest and typically occurs in museums, art galleries, and libraries. The task of organizing data around specific topics or themes is a vibrant and growing effort in the biological sciences but to date this effort has not been actively pursued in the Earth sciences. This presentation will introduce the concept of geocuration, which we define it as the act of searching, selecting, and synthesizing Earth science data/metadata and information from across disciplines and repositories into a single, cohesive, and useful compendium. We also present the Climate Data Initiative (CDI) project as an prototypical example. The CDI project is a systematic effort to manually curate and share openly available climate data from various federal agencies. CDI is a broad multi-agency effort of the U.S. government and seeks to leverage the extensive existing federal climate-relevant data to stimulate innovation and private-sector entrepreneurship to support national climate change preparedness. The geocuration process used in the CDI project, key lessons learned, and suggestions to improve similar geocuration efforts in the future will be part of this presentation.

climate

Governing Data Findability, Accessibility, Interoperability and Reusability (FAIR) Compliance

The most recent data strategy documents at both the federal and NASA levels stipulate that systems should strive for the data they manage to be Findable, Accessible, Interoperable, and Reusable (FAIR). The NASA Life Sciences Portal (NLSP) has already begun leading efforts in this area for HRP, initiating efforts to comply with the FAIR principles. The broad interpretation of the FAIR principles has led to a plethora of tools that use a splay of metrics specifically but variably developed to judge how compliant data and systems are with the principles. A recent review [3] identified and studied 1,180 metrics across 20 publicly available tools for checking FAIR compliance of data and systems. Because of their very recent development, many organizations and data systems managers and developers have not yet had adequate time or resources to understand these FAIR compliance tools and metrics, their variations in design, accuracy or ease of application to their specific data sets and systems. Thus, it would be best for larger organizations like NASA to approach formulating a strategy for governance of FAIR compliance that can be flexibly applied and is adaptable to an evolving awareness knowledge of FAIR compliance methods and tools. In September 2024, the NASA Science Mission Directorate(SMD) organized a workshop on NASA science data repositories, including the topics of implementing FAIR and governing FAIR compliance across SMD. The initial part of these FAIR discussions focused on developing consensus around required science metadata fields. This is challenging given the diverse nature of NASA’s scientific data portfolio, the variety of metadata models and vocabularies used, and variable level of resources available to curate these data. Later discussion focused on three possible approaches to governing FAIR compliance: distributed, in which various programs, projects or systems define their own methods for assessing FAIR compliance, reporting results up appropriate management lines; centralized, in which higher-level organization(s) specify compliance tools or methods for the various data systems; and multi-level, in which a group comprised of individuals with expertise from multiple levels with organizations is formed to provide guidance and/or specifications for governing FAIR compliance. We report on the recommendations this session yielded, and how these might be shaped specifically to help implement and govern the compliance with FAIR of Human Research Program data and systems.

governance

Introducing and Evaluating the Climate Hazards Center IMERG with Stations (CHIMES) Timely Station-Enhanced Integrated Multisatellite Retrievals for Global Precipitation Measurement

As human exposure to hydroclimatic extremes increase and the number of in situ precipitation observations declines, precipitation estimates, such as those provided by the Integrated Multisatellite Retrievals for Global Precipitation Measurement (GPM) (IMERG) mission, provide a critical source of information. Here, we present a new gauge-enhanced dataset [the Climate Hazards Center IMERG with Stations (CHIMES)] designed to support global crop and hydrologic modeling and monitoring. CHIMES enhances the IMERG Late Run product using an updated Climate Hazards Center (CHC) high-resolution climatology (CHPclim) and low-latency rain gauge observations. CHPclim differs from other products because it incorporates long-term averages of satellite precipitation, which increases CHPclim ’s fidelity in data-sparse areas with complex terrain. This fidelity translates into performance increases in unbiased IMERGlate data, which we refer to as CHIME. This is augmented with gauge observations to produce CHIMES. The CHC’s curated rain gauge archive contains valuable contributions from many countries. There are two versions of CHIMES: preliminary and final. The final product has more copious and better-curated station data. Every pentad and month, bias-adjusted IMERGlate fields are combined with gauge observations to create pentadal and monthly CHIMESprelim and CHIMESfinal. Comparisons with pentadal, high-quality gridded station data show that IMERG late performs well (r = 0.75), but has some systematic biases which can be reduced. Monthly cross-validation results indicate that unbiasing increases the variance explained from 50% to 63% and decreases the mean absolute error from 48 to 39 mm month −1. Gauge enhancement then increases the variance explained to 75%, reducing the mean absolute error to 27 mm month −1.

Chris C funk

Deploying Object Oriented Data Technology to the Planetary Data System

How do you provide more than 350 scientists and researchers access to data from every instrument in Odyssey when the data is curated across half a dozen institutions and in different formats and is too big to mail on a CD-ROM anymore? The Planetary Data System (PDS) faced this exact question. The solution was to use a metadata-based middleware framework developed by the Object Oriented Data Technology task at NASA s Jet Propulsion Laboratory. Using OODT, PDS provided - for the first time ever - data from all mission instruments through a single system immediately upon data delivery.

Kelly, S.

Open-Source Science-led Development of the Atmosphere Observing System (AOS) Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David Giles

Open-Source Science-led Development of the AOS Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David M. Giles

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data

Organic Contamination Baseline Study: In NASA JSC Astromaterials Curation Laboratories. Summary Report

In preparation for OSIRIS-REx and other future sample return missions concerned with analyzing organics, we conducted an Organic Contamination Baseline Study for JSC Curation Labsoratories in FY12. For FY12 testing, organic baseline study focused only on molecular organic contamination in JSC curation gloveboxes: presumably future collections (i.e. Lunar, Mars, asteroid missions) would use isolation containment systems over only cleanrooms for primary sample storage. This decision was made due to limit historical data on curation gloveboxes, limited IR&D funds and Genesis routinely monitors organics in their ISO class 4 cleanrooms.

Calaway, Michael J.