Search NASA⌕ Search

SEARCH · Search NASA

Results for “Crowdsourcing Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William↗

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

A Cloud-Based Global Flood Disaster Community Cyber-Infrastructure: Development and Demonstration

Flood disasters have significant impacts on the development of communities globally. This study describes a public cloud-based flood cyber-infrastructure (CyberFlood) that collects, organizes, visualizes, and manages several global flood databases for authorities and the public in real-time, providing location-based eventful visualization as well as statistical analysis and graphing capabilities. In order to expand and update the existing flood inventory, a crowdsourcing data collection methodology is employed for the public with smartphones or Internet to report new flood events, which is also intended to engage citizen-scientists so that they may become motivated and educated about the latest developments in satellite remote sensing and hydrologic modeling technologies. Our shared vision is to better serve the global water community with comprehensive flood information, aided by the state-of-the- art cloud computing and crowdsourcing technology. The CyberFlood presents an opportunity to eventually modernize the existing paradigm used to collect, manage, analyze, and visualize water-related disasters.

CyberFlood↗

The silicon citizen naturalist

Smartphone-wielding citizen scientists and an AI called FLORIST are transforming ecology at the continental scale. Here, in this issue of Cell, when Tibbs-Cortes et al. pair the crowdsourced data with controlled genetics, they discover how switchgrass times its flowering to outwit both frost and heat, depending on latitude.

Hudson, Matthew E. [University of Illinois at Urba↗

Automatic Lane-Level Road Network Extraction from Aerial Imagery for Transportation Digital Twins

Accurate road networks are essential for credible traffic microsimulation and transportation digital twins, yet high-definition maps are often difficult to obtain due to limited availability, high cost, or proprietary restrictions. Some build networks from crowdsourced data, such as OpenStreetMap, but these sources often contain geometric and semantic inconsistencies. Others create networks manually, a process that is labor-intensive and difficult to scale. To address these limitations, this work presents an end-to-end pipeline that automatically extracts georeferenced, lane-level road networks from publicly available high-resolution satellite imagery and converts them into simulation-ready assets. The developed end-to-end pipeline has three primary modules: (1) A computer-vision-based module first detects directed lane geometries and intersection layouts. (2) A heuristic-based topology construction module then identifies approach and exit legs and establishes conflict-free lane-to-lane connections. (3) Finally, an automatic simulation-building module converts the extracted network into standard formats, e.g., OpenDRIVE, and generates routable SUMO networks. The framework supports both complete network construction from scratch and local-scale refinement of existing networks through lane-count correction, transition recovery, and geometric regularization. The proposed pipeline provides a practical pathway to generate traffic simulation networks from satellite imagery, significantly reducing manual reconstruction effort and enabling scalable, continuously updated transportation digital twins.

Guo, Hetian [University of Georgia, Athens] (ORCID↗

Examining Rail Transportation Route of Crude Oil in the United States Using Crowdsourced Social Media Data

Safety issues associated with transporting crude oil by rail have been a concern since the boom of the U.S. domestic shale oil production in 2012. During the last decade, over 300 crude-oil-by-rail incidents have occurred in the United States. Some of them have caused adverse consequences including fire and hazardous materials leakage. However, only limited information on crude-on-rail routes and their associated risks is available to the public. To this end, this study proposed an unconventional way to reconstruct crude-on-rail routes using geotagged photos harvested from the Flickr website. The proposed method linked the geotagged photos of crude oil trains posted online with national railway networks to identify potential railway segments that those crude oil trains were traveling on. Here, a shortest path-based method was applied to infer the complete crude-on-rail routes, by utilizing the confirmed railway segments as well as their directional information. Validation of the inferred routes was performed using a public map and official crude oil incident data. The results suggested that the inferred routes based on geotagged photos had high coverage, with approximately 96% of the documented crude oil incidents aligned with the reconstructed crude-on-rail network. The inferred crude oil train routes were found to pass through several metropolitan areas of high population density, who were exposed to potential risk. These findings could improve situational awareness for policy makers and transportation planners. In addition, with the inferred routes, this study has established a good foundation for future crude oil train risk-analyses along the rail route.

42 ENGINEERING↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

Urban Versus Lake Impacts on Heat Stress and Its Disparities in a Shoreline City

Abstract Shoreline cities are influenced by both urban‐scale processes and land‐water interactions, with consequences on heat exposure and its disparities. Heat exposure studies over these cities have focused on air and skin temperature, even though moisture advection from water bodies can also modulate heat stress. Here, using an ensemble of model simulations covering Chicago, we find that Lake Michigan strongly reduces heat exposure (2.75°C reduction in maximum average air temperature in Chicago) and heat stress (maximum average wet bulb globe temperature reduced by 0.86°C) during the day, while urbanization enhances them at night (2.75 and 1.57°C increases in minimum average air and wet bulb globe temperature, respectively). We also demonstrate that urban and lake impacts on temperature (particularly skin temperature), including their extremes, and lake‐to‐land gradients, are stronger than the corresponding impacts on heat stress, partly due to humidity‐related feedback. Likewise, environmental disparities across community areas in Chicago seen for skin temperature are much higher (1.29°C increase for maximum average values per $10,000 higher median income per capita) than disparities in air temperature (0.50°C increase) and wet bulb globe temperature (0.23°C increase). The results call for consistent use of physiologically relevant heat exposure metrics to accurately capture the public health implications of urbanization.

54 ENVIRONMENTAL SCIENCES↗

Linkages Between Mineral Element Composition of Soils and Sediments With Hyporheic Zone Dissolved Organic Matter Chemistry Across the Contiguous United States

The hyporheic zone is a hotspot for biogeochemical cycling where interactions with mineral metals preserve the release and biodegradation of organic matter (OM). A small fraction of OM can still be exchanged between localized sediments and the overlying water column, and recent evidence suggests there exists a longitudinal structuring in sediment dissolved OM (DOM) chemistry across the continental United States (CONUS). In this study, we tested a hypothesis that water extractable sediment DOM chemistry could be explained by sediment metal contents and integrative watershed scale features at the CONUS scale. Crowdsourced samples were characterized for high resolution mass spectrometry and coupled with sediment metals determined via x-ray fluorescence as well as with land cover and soil elemental information obtained from national databases. Our results highlight weak relationships between DOM chemistry and elemental composition at the CONUS scale indicating limited transferability of organo-metal linkages into multi-scale hydrobiogeochemical models.

58 GEOSCIENCES↗

Labeling sequential data from noisy annotations

Crowdsourcing algorithms often work under the assumption that the data samples are independent. Recent work has shown that data dependence, such as temporal correlations in sequential data, can be leveraged to improve the label quality. Existing methods that exploit this special structure rely on third-order statistics of the annotator outputs to ensure the identifiability of key latent parameters, which are costly to acquire. This work proposes an approach for integrating crowdsourced annotations under the Dawid-Skene/Hidden Markov Model (DS-HMM) for sequential data based on second-order statistics, which naturally enjoys a lower sample complexity. An effective algorithm is proposed to tackle the challenging optimization problem associated with the proposed estimator. Numerical experiments showcase the effectiveness of the data labeling paradigm.

Marrinan, Timothy P.↗

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif↗

Estimating building occupancy: a machine learning system for day, night, and episodic events

Building occupancy research increasingly emphasizes understanding the social and physical dynamics of how people occupy space. Opportunities in the open source domain including social media, Volunteered Geographic Information, crowdsourcing, and sensor data have proliferated, resulting in the exploration of building occupancy dynamics at varying spatiotemporal scales. At Oak Ridge National Laboratory, research into building occupancies through the development of a global learning framework that accommodates exploitation of open source authoritative sources, including governmental census and surveys, journal articles, real estate databases, and more, to report national and subnational building occupancies across the world continues through the Population Density Tables (PDT) project. This probabilistic learning system accommodates expert knowledge, experience, and open-source data to capture local, socioeconomic, and cultural information about human activity. It does so through a systematic process of data harmonization techniques in the development of observation models for over 50 building types to dynamically update baseline estimates and report probabilistic diurnal and episodic building occupancy estimates. This discussion will explore how PDT is implemented at scale and expanded based on the development of observation model classes and will explain how to interpret and spatially apply the reported probability occupancy estimates and uncertainty.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Determining the biogeochemical transformations of organic matter composition in rivers using molecular signatures

Inland waters are hotspots for biogeochemical activity, but the environmental and biological factors that govern the transformation of organic matter (OM) flowing through them are still poorly constrained. Here we evaluate data from a crowdsourced sampling campaign led by the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) consortium to investigate broad continental-scale trends in OM composition compared to localized events that influence biogeochemical transformations. Samples from two different OM compartments, sediments and surface water, were collected from 97 streams throughout the Northern Hemisphere and analyzed to identify differences in biogeochemical processes involved in OM transformations. By using dimensional reduction techniques, we identified that putative biogeochemical transformations and microbial respiration rates vary across sediment and surface water along river continua independent of latitude (18°N–68°N). In contrast, we reveal small- and large-scale patterns in OM composition related to local (sediment vs. water column) and reach (stream order, latitude) characteristics. These patterns lay the foundation to modeling the linkage between ecological processes and biogeochemical signals. We further showed how spatial, physical, and biogeochemical factors influence the reactivity of the two OM pools in local reaches yet find emergent broad-scale patterns between OM concentrations and stream order. OM processing will likely change as hydrologic flow regimes shift and vertical mixing occurs on different spatial and temporal scales. As our planet continues to warm and the timing and magnitude of surface and subsurface flows shift, understanding changes in OM cycling across hydrologic systems is critical, given the unknown broad-scale responses and consequences for riverine OM.

rivers↗

Integrating Machine Learning into a Crowdsourced Model for Earthquake-Induced Damage Assessment

On January 12th, 2010, a catastrophic 7.0M earthquake devastated the country of Haiti. In the aftermath of an earthquake, it is important to rapidly assess damaged areas in order to mobilize the appropriate resources. The Haiti damage assessment effort introduced a promising model that uses crowdsourcing to map damaged areas in freely available remotely-sensed data. This paper proposes the application of machine learning methods to improve this model. Specifically, we apply work on learning from multiple, imperfect experts to the assessment of volunteer reliability, and propose the use of image segmentation to automate the detection of damaged areas. We wrap both tasks in an active learning framework in order to shift volunteer effort from mapping a full catalog of images to the generation of high-quality training data. We hypothesize that the integration of machine learning into this model improves its reliability, maintains the speed of damage assessment, and allows the model to scale to higher data volumes.

crowdsourcing↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗