Search NASA⌕ Search

SEARCH · Search NASA

Results for “citizen science data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Perspectives on Citizen Science Data Quality

Information about data quality helps potential data users to determine whether and how data can be used and enables the analysis and interpretation of such data. Providing data quality information improves opportunities for data reuse by increasing the trustworthiness of the data. Recognizing the need for improving the quality of citizen science data, we describe quality assessment and quality control (QA/QC) issues for these data and offer perspectives on aspects of improving or ensuring citizen science data quality and for conducting research on related issues.

54 ENVIRONMENTAL SCIENCES↗

NASA ESDS Citizen Science Data Working Group

This document provides guidelines for legal, policy, and ethical issues; standards for citizen science data collection and management; information on ensuring usability of citizen science data and communication regarding its use; and best practices for long-term archival of citizen science data.Section1 contains a detailed discussion of policy, ethical, and legal considerations influencing citizen science data collection. Section 2 considers standards for documentation, including documentation of instrumentation, procedures, and the data itself. It concludes with a discussion of how citizen science data should be attributed. Section 3 provides guidance about how to ensure citizen science data are collected and stored in a useable way. It also considers how NASA and data producers should notify the scientific community, including citizen scientists and the public, about citizen science datasets and the scientific conclusions reached using them. Finally, Section 4 provides detailed information regarding what should be archived from projects using a citizen science approach, including data and code. It provides guidance about archive location, process, and timeframe, as well as information about data access and distribution services provided by NASA that may be relevant to data producers working with citizen scientists.

Citizen Science↗

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data↗

Clouds Around the World: How a Simple Citizen Science Data Challenge Became a Worldwide Success

Citizen science is often recognized for its potential to directly engage the public in science, and is uniquely positioned to support and extend participants’ learning in science. In March 2018, the Global Learning and Observations to Benefit the Environment (GLOBE) Program, NASA’s largest and longest-lasting citizen science program about Earth, organized a month-long event that asked people around the world to contribute daily cloud observations and photographs of the sky (15 March–15 April 2018). What was considered a simple engagement activity turned into an unprecedented worldwide event that garnered major public interest and media recognition, collecting over 55,000 observations from 99 different countries, in more than 15,000 locations, on every continent including Antarctica. The event was called the “Spring Cloud Challenge” and was created to 1) engage the general public in the scientific process and promote the use of the GLOBE Observer app, 2) collect ground-based visual observations of varying cloud types during boreal spring, and 3) increase the number and locations of ground-based visual cloud observations collocated with cloud-observing satellites. The event resulted in roughly 3 times more observations than during the historic and highly publicized 2017 North American total solar eclipse. The dataset also includes observations over the Drake Passage in Antarctica and reports from intense Saharan dust events. This article describes how the challenge was crafted, outreach to volunteer scientists around the world, details of the data collected, and impact of the data.

GLOBE↗

Leveraging Machine Learning and Geo-Tagged Citizen Science Data to Disentangle the Factors of Avian Mortality Events at the Species Level

Abrupt environmental changes can affect the population structures of living species and cause habitat loss and fragmentations in the ecosystem. During August–October 2020, remarkably high mortality events of avian species were reported across the western and central United States, likely resulting from winter storms and wildfires. However, the differences of mortality events among various species responding to the abrupt environmental changes remain poorly understood. In this study, we focused on three species, Wilson’s Warbler, Barn Owl, and Common Murre, with the highest mortality events that had been recorded by citizen scientists. We leveraged the citizen science data and multiple remotely sensed earth observations and employed the ensemble random forest models to disentangle the species responses to winter storm and wildfire. We found that the mortality events of Wilson’s Warbler were primarily impacted by early winter storms, with more deaths identified in areas with a higher average daily snow cover. The Barn Owl’s mortalities were more identified in places with severe wildfire-induced air pollution. Both winter storms and wildfire had relatively mild effects on the mortality of Common Murre, which might be more related to anomalously warm water. Our findings highlight the species-specific responses to environmental changes, which can provide significant insights into the resilience of ecosystems to environmental change and avian conservations. Additionally, the study emphasized the efficiency and effectiveness of monitoring large-scale abrupt environmental changes and conservation using remotely sensed and citizen science data.

47 OTHER INSTRUMENTATION↗

Assimilation of citizen science data in snowpack modeling using a new snow data set: Community Snow Observations

A physically based snowpack evolution and redistribution model was used to test the effectiveness of assimilating crowd-sourced snow depth measurements collected by citizen scientists. The Community Snow Observations project gathers, stores, and distributes measurements of snow depth recorded by recreational users and snow professionals in high mountain environments. These citizen science measurements are valuable since they come from terrain that is relatively undersampled and can offer in situ snow information in locations where snow information is sparse or nonexistent. The present study investigates (1) the improvements to model performance when citizen science measurements are assimilated, and (2) the number of measurements necessary to obtain those improvements. Model performance is assessed by comparing time series of observed (snow pillow) and modeled snow water equivalent values, by comparing spatially distributed maps of observed (remotely sensed) and modeled snow depth, and by comparing fieldwork results from within the study area. The results demonstrate that few citizen science measurements are needed to obtain improvements in model performance, and these improvements are found in 62 % to 78 % of the ensemble simulations, depending on the model year. Model estimations of total water volume from a subregion of the study area also demonstrate improvements in accuracy after CSO measurements have been assimilated. These results suggest that even modest measurement efforts by citizen scientists have the potential to improve efforts to model snowpack processes in high mountain environments, with implications for water resource management and process-based snow modeling.

54 ENVIRONMENTAL SCIENCES↗

Citizen Science Data Quality: The GLOBE Program

The Global Learning and Observations to Benefit the Environment (GLOBE) Program is an international program that provides a way for students and the public to contribute Earth system observations. Currently 122 countries, more than 40,000 schools, and 200,000 citizen scientists are participating in GLOBE. Since 1995, participants have contributed 195 million observations. Modes of data collection and data entry have evolved with technology over the lifetime of the program, including the launch of the GLOBE Observer mobile app in 2016 to broaden access and public participation in data collection. GLOBE must meet the data needs of a diverse range of stakeholders, from elementary school classrooms to scientists across the globe, including NASA scientists. Operational quality assurance measures include participant training, adherence to standardized data collection protocols, range and logic checks, and an approval process for photos submitted with an observation. In this presentation, we will discuss the current state of operational data QA/QC, as well as additional QA/QC processes recently explored and future directions.

Amos, Helen↗

A Case Study Comparing Citizen Science Aurora Data with Global Auroral Boundaries Derived from Satellite Imagery and Empirical Models

Aurorasaurus is a citizen science project that offers a new, global data source consisting of ground-based reports of the aurora. For this case study, aurora data collected during the 17-18 March 2015 geomagnetic storm are examined to identify their conjunctions with Defense Meteorological Satellite Program (DMSP) satellite passes over the high latitude auroral regions. This unique set of aurora data can provide ground-truth validation of existing auroral precipitation models. Particularly, the solar wind driven, Oval Variation, Assessment, Tracking, Intensity, and Online Nowcasting (OVATION) Prime 2013 (OP-13) model and a Kp-dependent model of Zhang-Paxton (Z-P) are utilized for our boundary validation efforts. These two similar models are compared for the first time. Global equatorward auroral boundaries are derived from the OP 13 model and the DMSP Special Sensor Ultraviolet Spectrographic Imager (SSUSI) far ultraviolet (FUV) data using the Z-P model at a fixed flux level of 0.2 erg cm(exp -2)s(exp -1). These boundaries are then compared with citizen science reports as well as with each other. Even though there are some large differences between the global boundaries for a few cases, the average difference is about 1.5 deg in geomagnetic latitude, with OP-13 being equatorward of Z-P model. When these boundaries are compared with each other as a function of local time, no clear overall trend as a function of local time was observed. It is also found that the ground based reports are more consistent with the predictions of the OP-13 model.

Kosar, Burcu C.↗

Aurorasaurus Database of Real-Time, Crowd-Sourced Aurora Data for Space Weather Research

This technical report documents the details of Aurorasaurus citizen science data for the period spanning 2015 and 2016 as well as its routine data filtering protocols. Aurorasaurus citizen science data is a collection of auroral sightings submitted to the project via its website or apps and mined from social media. It is a robust data set and particularly abundant during strong geomagnetic storms when auroral precipitation models have the highest uncertainty. These data are offered to the scientific community for use through an openaccess database in its raw and scientific formats, each of which is described in detail in this technical report. Furthermore, by demonstrating its scientific utility, we aim to encourage its integration into auroral research.

Citizen science↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

Citizen Science Twitter Data Management for Earth Science Applications

Social media data can provide useful real-time and historical information relating to the natural world, but managing this data poses challenges. Scientists at GES DISC are exploring the potential of Twitter data to augment precipitation data from the Global Precipitation Measurement (GPM) mission. However, the format of Twitter data is unconventional in the context of NASA data centers, resulting in frustration for scientists who need to work with the data. This study investigated procedures and standards needed to properly manage Twitter data to make them compatible with these data centers. After comparing databases, the study found that the MongoDB database was best suited for the storage of raw Twitter data due to its flexibility, ability to be accessed by multiple users, and querying functionality. The study used the Python package Zarr to transform processed Twitter data into a gridded format similar to that of satellite data. Each Tweet was mapped onto a time-space grid; each grid location contained information about Tweet attributes and precipitation. The study developed a pipeline for downloading, storing, and gridding Twitter data and transformed Twitter data into an understandable format for users of NASA satellite data.

Li, Rachel↗

NASA GLOBE CLOUD GAZE: Creating Data Quality Flags for Citizen Science Cloud Observations Matched to NASA Satellite Data

The GLOBE Program, NASA’s largest and longest lasting citizen science program about the Earth, has been collecting cloud observations matched to multiple satellite data daily. The program’s cloud protocol is historically the most popular protocol as your eyes are the only instruments you need to collect observations of the sky. This dataset includes over 3,300,000 cloud observations with variables like total cloud cover, cloud type and opacity that are collocated to the nearest overpass times of geostationary satellites (GOES-15, GOES-16, GOES-17, Meteosat-8, Meteosat-11, or Himawari-8), or to Clouds and the Earth’s Radiant Energy System (CERES) instruments onboard Aqua and Terra, or the Cloud–Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) satellite. In order to increase the usability of this dataset, the Community science project Leveraging Online and User Data through GLOBE And Zooniverse Engagement (CLOUD GAZE) has been developed to generate data quality flags of these ground-up and top-down perspectives of sky and clouds. Recently funded through NASA’s Citizen Science for Earth Systems Program, CLOUD GAZE has partnered with the Zooniverse online platform to obtain reference data and image tagging of sky photographs collected through The GLOBE Program’s clouds protocol. This paper will present the GLOBE Clouds dataset matched to NASA satellite data, integration of CLOUD GAZE to develop data quality flags, and research applications of the dataset (includes ground-up and top-down perspective comparisons, ground observations of dust storms and smoke plumes, and cloud observations in polar regions). The paper will also present on techniques and recommendations for classroom use and for community engagement, particularly for those looking to online resources.

Marilé Colón Robles↗

Building a framework to genetically characterize “feather spots” and understand demographic impacts of solar energy sites on migratory bird populations

The lack of data on the impact of utility-scale solar facilities on avian species and populations adds to the cost of siting and operation. As much as 32 percent of the avian biological material (feathers and carcasses) recovered from solar facilities remain unidentified, because they often take the form of “feather spots”. Feather spots are remains of impacted animals that can be separated into two broad categories: 1) those remains that may be visually identified to a species, or 2) those that cannot be visually identified to a species due to degradation from the environment and/or scavenger activity (listed as “unknown”). Even when feather spots can be identified to species, they cannot be visually assigned to particular breeding populations. In some cases, it is unknown whether multiple feather spots represent single or multiple individuals. This project’s objectives were to: 1. Use a developed, genetic-based technique to identify and determine the species, population of origin, and number of individuals found in feather spots recovered from solar facilities. 2. Implement collected data and resulting analyses to develop a publicly accessible web-based decision-making tool that can be used by the solar industry, regulators and other stakeholders to inform siting, mitigation, and conservation management efforts. 3. Establish a not-for-profit fee-for-service center at UCLA to ensure collection and identification of feather spots continue after the project period of performance. During the Project Period, we proposed to establish a pipeline for collecting, transporting, and storing of avian biological material collected at solar facilities and the collection and identification of feather spots to species and individual. We proposed the development of a genetic-based framework that would recover viable DNA from feather spots, amplify this DNA (i.e., make millions of copies of the original DNA), and use it to match the resulting sequences to a national database of known species of birds. The result would be the identification of feathers spots that were previously unidentified, and the incorporation of these samples into a larger database that included all samples recovered from solar facilities. The resulting report (below) details the result of this work and its alignment with proposed activities. We proposed the use of the data collected to assess the comparative risk to specific species or populations of species from solar facilities. For some species, we have already identified genomic markers of specific breeding populations and developed “genoscapes,” maps of unique genetic variation across the full breeding range of a species. We used these (previously and newly developed) genoscapes to probabilistically link a feather spot to the specific breeding populations from which it originated (assignment probabilities range from 75%-100% depending on species and population groups). For those species without genoscapes, we developed a vulnerability and susceptibility estimate that determines the relative local and regional risk to populations that are in geographic proximity to solar facilities, using citizen science data (Breeding Bird Survey (BBS) and eBird). These two feather spot processing pipelines (see Figure 1 below) provide quantitative estimates as to the numbers of individuals from a given population of origin that are affected by solar facilities, and ultimately can reduce costs to the consumer by reducing the industry costs associated with mitigation and siting strategies for future solar energy development.

14 SOLAR ENERGY↗

From Toes to Top-of-the-Atmosphere: Fowler Sneaker Index

Fowler Sneaker Index (FSI), developed by a NASA summer intern, is a new Ocean Color application that facilitates continuous monitoring of environmental conditions in the Chesapeake Bay. It builds on three decades of citizen science data collected by former Maryland State Senator Bernie Fowler, during his yearly "Wade-ins in the Patuxent River". FSI demonstrates how NASA's Earth-observing tools, in combination with a concerned and engaged public, can take science from the tips of our toes-to-top-of the atmosphere and back.

Optics Express↗

Do Citizen Science Intense Observation Periods Increase Data Usability? A Deep Dive of the NASA GLOBE Clouds Data Set With Satellite Comparisons

The Global Learning and Observations to Benefit the Environment (GLOBE) citizen science program has recently conducted a series of month-long intensive observation periods (IOPs), asking the public to submit daily reports on cloud and sky conditions from all regions of Earth. This provides a wealth of crowdsourced observations from the ground, which complements other conventional scientific cloud data. In addition, the GLOBE reports are matched in space and time with geostationary and low Earth orbit satellites, which allows for a straightforward comparison of cloud properties, and minimizes the biases associated with mismatched sampling between participants and satellites. The matched GLOBE dataset is used to calculate the mean observed cloud cover by atmospheric level both worldwide and by region. The overall magnitudes of cloud cover between the GLOBE participants and the matched satellites agree within 10%, which is notable given the distinctly different natures of the data sources. The mean vertical cloud profiles show GLOBE reporting more low-level clouds and fewer high-level clouds than satellites. The low cloud disagreement is likely related to satellites missing low clouds when high clouds block their view. Conversely, the high cloud disagreement is related primarily to cloud opacity, as satellites may miss some optically thin clouds. Monte Carlo testing shows the results to be robust, and the tripled amount of IOP data reduces uncertainty by half. These findings also highlight ways in which citizen science IOP data may be used to support scientific research while accounting for their unique properties. Plain Language Summary: Citizen science is becoming an increasingly prominent aspect of scientific research, and so it important to study how citizen science data can be used effectively. For example, The GLOBE Program has recently conducted a series of special data-collecting events, or “challenges”, which gathered large numbers of reports on cloud and sky conditions. Because NASA GLOBE Clouds matches the participant reports with cloud observations from satellites, we can use these data to get a combined view of clouds from above and below. When looking at the average cloud cover for different atmospheric levels across Earth, we find that the GLOBE participants and the satellites agree quite closely. This is a surprising and fascinating find, given how different in nature volunteer ground reports are to satellite measurements. However, there are some small but notable disagreements between GLOBE participants and satellites about the distribution of cloud cover at different levels. In addition, by testing the data for uncertainty, we show that the results from the GLOBE data are reliable, and that more public participation improves the reliability. So, by carefully designing the analysis methodology, and by testing for the uncertainty of the data, citizen science can make a meaningful contribution to scientific research.

J. Brant Dodson↗

GLOBE Observer and the GO on a Trail Data Challenge: A Citizen Science Approach to Generating a Global Land Cover Land Use Reference Dataset

Land cover and land use are highly visible indicators of climate change and human disruption to natural processes. While land cover is frequently monitored over a large area using satellite data, ground-based reference data is valuable as a comparison point. The NASA-funded GLOBE Observer (GO) program provides volunteer-collected land cover photos tagged with location, date and time, and, in some cases, land cover type. When making a full land cover observation, volunteers take six photos of the site, one facing north, south, east, and west (N-S-E-W) respectively, one pointing straight up to capture canopy and sky, and one pointing down to document ground cover. Together, the photos document a 100 meter square of land. Volunteers may then optionally tag each N-S-E-W photo with the land cover types present. Volunteers collect the data through a smartphone app, also called GLOBE Observer, resulting in consistent data. While land cover data collected through GLOBE Observer is ongoing, this paper presents the results of a data challenge held between June 1 and October 15, 2019. Called “GO on a Trail,” the challenge focused on collecting land cover observations at designated points along the Lewis and Clark National Historic Trail (LCNHT), which transects approximately 7,900 kilometers of the United States from east to west. The challenge also included an international component in which participants were recognized for the most data collected during the period. The challenge resulted in more than 2800 land cover data points from around the world with just over 900 land cover data points along the LCNHT. The data collected along the LCNHT is largely within a 1,000 meter corridor along the historic route. Internationally, the challenge had strong participation in Australia, where a Scouts Australia team-based competition was conducted in conjunction with GO on a Trail. GLOBE Observer collections can serve as reference data, ground truthing satellite imagery for the improvement and verification of broad land cover maps. Continued collection using this protocol will build a database documenting climate-related land cover and land use change into the future.

land cover↗