Search NASASearch

SEARCH · Search NASA

Results for “Data usability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Do Citizen Science Intense Observation Periods Increase Data Usability? A Deep Dive of the NASA GLOBE Clouds Data Set With Satellite Comparisons

The Global Learning and Observations to Benefit the Environment (GLOBE) citizen science program has recently conducted a series of month-long intensive observation periods (IOPs), asking the public to submit daily reports on cloud and sky conditions from all regions of Earth. This provides a wealth of crowdsourced observations from the ground, which complements other conventional scientific cloud data. In addition, the GLOBE reports are matched in space and time with geostationary and low Earth orbit satellites, which allows for a straightforward comparison of cloud properties, and minimizes the biases associated with mismatched sampling between participants and satellites. The matched GLOBE dataset is used to calculate the mean observed cloud cover by atmospheric level both worldwide and by region. The overall magnitudes of cloud cover between the GLOBE participants and the matched satellites agree within 10%, which is notable given the distinctly different natures of the data sources. The mean vertical cloud profiles show GLOBE reporting more low-level clouds and fewer high-level clouds than satellites. The low cloud disagreement is likely related to satellites missing low clouds when high clouds block their view. Conversely, the high cloud disagreement is related primarily to cloud opacity, as satellites may miss some optically thin clouds. Monte Carlo testing shows the results to be robust, and the tripled amount of IOP data reduces uncertainty by half. These findings also highlight ways in which citizen science IOP data may be used to support scientific research while accounting for their unique properties. Plain Language Summary: Citizen science is becoming an increasingly prominent aspect of scientific research, and so it important to study how citizen science data can be used effectively. For example, The GLOBE Program has recently conducted a series of special data-collecting events, or “challenges”, which gathered large numbers of reports on cloud and sky conditions. Because NASA GLOBE Clouds matches the participant reports with cloud observations from satellites, we can use these data to get a combined view of clouds from above and below. When looking at the average cloud cover for different atmospheric levels across Earth, we find that the GLOBE participants and the satellites agree quite closely. This is a surprising and fascinating find, given how different in nature volunteer ground reports are to satellite measurements. However, there are some small but notable disagreements between GLOBE participants and satellites about the distribution of cloud cover at different levels. In addition, by testing the data for uncertainty, we show that the results from the GLOBE data are reliable, and that more public participation improves the reliability. So, by carefully designing the analysis methodology, and by testing for the uncertainty of the data, citizen science can make a meaningful contribution to scientific research.

J. Brant Dodson

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data

Advancing Open Science in Atmospheric Research: Integrating Data Usability and Machine Learning

In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences." "In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences.

Jennifer Wei

Large Scale Data Mining to Improve Usability of Data: An Intelligent Archive Testbed

Research in certain scientific disciplines - including Earth science, particle physics, and astrophysics - continually faces the challenge that the volume of data needed to perform valid scientific research can at times overwhelm even a sizable research community. The desire to improve utilization of this data gave rise to the Intelligent Archives project, which seeks to make data archives active participants in a knowledge building system capable of discovering events or patterns that represent new information or knowledge. Data mining can automatically discover patterns and events, but it is generally viewed as unsuited for large-scale use in disciplines like Earth science that routinely involve very high data volumes. Dozens of research projects have shown promising uses of data mining in Earth science, but all of these are based on experiments with data subsets of a few gigabytes or less, rather than the terabytes or petabytes typically encountered in operational systems. To bridge this gap, the Intelligent Archives project is establishing a testbed with the goal of demonstrating the use of data mining techniques in an operationally-relevant environment. This paper discusses the goals of the testbed and the design choices surrounding critical issues that arose during testbed implementation.

Ramapriyan, Hampapuram

Pilot climate data system

A usable data base, the Pilot climate Data System (PCDS) is described. The PCDS is designed to be an interactive, easy-to-use, on-line generalized scientific information system. It efficiently provides uniform data catalogs; inventories, and access method, as well as manipulation and display tools for a large assortment of Earth, ocean and atmospheric data for the climate-related research community. Researchers can employ the PCDS to scan, manipulate, compare, display, and study climate parameters from diverse data sets. Software features, and applications of the PCDS are highlighted.

Source record

Gravity Wave Energetics Determined From Coincident Space-Based and Ground-Based Observations of Airglow Emissions

Significant progress was made toward the goals of this proposal in a number of areas during the covered period. Section 5.1 contains a copy of the originally proposed schedule. The tasks listed below have been accomplished: (1) Construction of space-based observing geometry gravity wave model. This model has been described in detail in the paper accompanying this report (Section 5.2). It can simulate the observing geometry of both ground-based, and orbital instruments allowing comparisons to be made between them. (2) Comparisons of relative emission intensity, temperatures, and Krassovsky's ratio for space- and ground-based observing geometries. These quantities are used in gravity wave literature to describe the effects of the waves on the airglow. (3) Rejection of Bates [1992], and Copeland [1994] chemistries for gravity wave modeling purposes. Excessive 02(A(sup 13)(Delta)) production led to overproduction of O2(b(sup 1)(Sigma)), the state responsible for the emission of O2. Atmospheric band. Attempts were made to correct for this behavior, but could not adequately compensate for this. (4) Rejection of MSX dataset due to lack of coincident data, and resolution necessary to characterize the waves. A careful search to identify coincident data revealed only four instances, with only one of those providing usable data. Two high latitude overpasses and were contaminated by auroral emissions. Of the remaining two mid-latitude coincidences, one overflight was obscured by cloud, leaving only one ten minute segment of usable data. Aside from the statistical difficulties involved in comparing measurements taken in this short period, the instrument lacks the necessary resolution to determine the vertical wavelength of the gravity wave. This means that the wave cannot be uniquely characterized from space with this dataset. Since no observed wave can be uniquely identified, model comparisons are not possible.

Source record

Experiences in Bridging the Gap between Science and Decision Making at NASA's GSFC Earth Science Data and Information Services Center (GES DISC)

Recognizing the significance of NASA remote sensing Earth science data in monitoring and better understanding our planet s natural environment, NASA has implemented the Decision Support Through Earth Science Research Results program (NASA ROSES solicitations). a) This successful program has yielded several monitoring, surveillance, and decision support systems through collaborations with benefiting organizations. b) The Goddard Space Flight Center (GSFC) Earth Sciences Data and Information Services Center (GES DISC) has participated in this program on two projects (one complete, one ongoing), and has had opportune ad hoc collaborations gaining much experience in the formulation, management, development, and implementation of decision support systems utilizing NASA Earth science data. c) In addition, GES DISC s understanding of Earth science missions and resulting data and information, including data structures, data usability and interpretation, data interoperability, and information management systems, enables the GES DISC to identify challenges that come with bringing science data to decision makers. d) The purpose of this presentation is to share GES DISC decision support system project experiences in regards to system sustainability, required data quality (versus timeliness), data provider understanding of how decisions are made, and the data receivers willingness to use new types of information to make decisions, as well as other topics. In addition, defining metrics that really evaluate success will be exemplified.

Kempler, Steven

Data simulation for the Lightning Imaging Sensor (LIS)

This project aims to build a data analysis system that will utilize existing video tape scenes of lightning as viewed from space. The resultant data will be used for the design and development of the Lightning Imaging Sensor (LIS) software and algorithm analysis. The desire for statistically significant metrics implies that a large data set needs to be analyzed. Before 1990 the quality and quantity of video was insufficient to build a usable data set. At this point in time, there is usable data from missions STS-34, STS-32, STS-31, STS-41, STS-37, and STS-39. During the summer of 1990, a manual analysis system was developed to demonstrate that the video analysis is feasible and to identify techniques to deduce information that was not directly available. Because the closed circuit television system used on the space shuttle was intended for documentary TV, the current value of the camera focal length and pointing orientation, which are needed for photoanalysis, are not included in the system data. A large effort was needed to discover ancillary data sources as well as develop indirect methods to estimate the necessary parameters. Any data system coping with full motion video faces an enormous bottleneck produced by the large data production rate and the need to move and store the digitized images. The manual system bypassed the video digitizing bottleneck by using a genlock to superimpose pixel coordinates on full motion video. Because the data set had to be obtained point by point by a human operating a computer mouse, the data output rate was small. The loan and subsequent acquisition of a Abekas digital frame store with a real time digitizer moved the bottleneck from data acquisition to a problem of data transfer and storage. The semi-automated analysis procedure was developed using existing equipment and is described. A fully automated system is described in the hope that the components may come on the market at reasonable prices in the next few years.

Boeck, William L.

Rocket/Nimbus Sounder Comparison (RNSC)

The experimental results for radiance and temperature differences in the Wallops Island comparisons indicate that the differences between satellite and rocket systems are of the same order of magnitude as the differences among the various satellite and rocket sounders. The Arcasondes produced usable data to about 50 km, while the Datasondes require design modification. The SIRS and IRIS soundings provided usable data to 30 mb; extension of these soundings was also investigated.

Source record

Hierarchical Data Format for Earth Observing System Data Product Developer's Guide

The "Hierarchical Data Format for Earth Observing System" talk will address the best practices for creating ESDIS data products. The work presented is done in support of Data Product Developers Guide Working Group with mission "to help data product developers make data usable for end users". During the presentation, we will use some examples of NASA data products and show how to modify them to make data more usable.

Data usability

The integration of a LANDSAT analysis capability with a geographic information system

The integration of LANDSAT data was achieved through the development of a flexible, compatible analysis tool and using an existing data base to select the usable data from a LANDSAT analysis. The software package allows manipulation of grid cell data plus the flexibility to allow the user to include FORTRAN statements for special functions. Using this combination of capabilities the user can classify a LANDSAT image and then selectivity merge the results with other data that may exist for the study area.

Nordstrand, E. A.

Qualitative Analysis for Maintenance Process Assessment

In order to improve software maintenance processes, we first need to be able to characterize and assess them. These tasks must be performed in depth and with objectivity since the problems are complex. One approach is to set up a measurement-based software process improvement program specifically aimed at maintenance. However, establishing a measurement program requires that one understands the problems to be addressed by the measurement program and is able to characterize the maintenance environment and processes in order to collect suitable and cost-effective data. Also, enacting such a program and getting usable data sets takes time. A short term substitute is therefore needed. We propose in this paper a characterization process aimed specifically at maintenance and based on a general qualitative analysis methodology. This process is rigorously defined in order to be repeatable and usable by people who are not acquainted with such analysis procedures. A basic feature of our approach is that actual implemented software changes are analyzed in order to understand the flaws in the maintenance process. Guidelines are provided and a case study is shown that demonstrates the usefulness of the approach.

Brand, Lionel

Langley advanced real-time simulation (ARTS) system

A system of high-speed digital data networks was developed and installed to support real-time flight simulation at the NASA Langley Research Center. This system, unlike its predecessor, employs intelligence at each network node and uses distributed 10-V signal conversion equipment rather than centralized 100-V equipment. A network switch, which replaces an elaborate system of patch panels, allows the researcher to construct a customized network from the 25 available simulation sites by invoking a computer control statement. The intent of this paper is to provide a coherent functional description of the system. This development required many significant innovations to enhance performance and functionality such as the real-time clock, the network switch, and improvements to the CAMAC network to increase both distances to sites and data rates. The system has been successfully tested at a usable data rate of 24 M. The fiber optic lines allow distances of approximately 1.5 miles from switch to site. Unlike other local networks, CAMAC does not buffer data in blocks. Therefore, time delays in the network are kept below 10 microsec total. This system underwent months of testing and was put into full service in July 1987.

Crawford, Daniel J.