Search NASA⌕ Search

SEARCH · Search NASA

Results for “Crowdsourcing Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

The silicon citizen naturalist

Smartphone-wielding citizen scientists and an AI called FLORIST are transforming ecology at the continental scale. Here, in this issue of Cell, when Tibbs-Cortes et al. pair the crowdsourced data with controlled genetics, they discover how switchgrass times its flowering to outwit both frost and heat, depending on latitude.

Hudson, Matthew E. [University of Illinois at Urba↗

Automatic Lane-Level Road Network Extraction from Aerial Imagery for Transportation Digital Twins

Accurate road networks are essential for credible traffic microsimulation and transportation digital twins, yet high-definition maps are often difficult to obtain due to limited availability, high cost, or proprietary restrictions. Some build networks from crowdsourced data, such as OpenStreetMap, but these sources often contain geometric and semantic inconsistencies. Others create networks manually, a process that is labor-intensive and difficult to scale. To address these limitations, this work presents an end-to-end pipeline that automatically extracts georeferenced, lane-level road networks from publicly available high-resolution satellite imagery and converts them into simulation-ready assets. The developed end-to-end pipeline has three primary modules: (1) A computer-vision-based module first detects directed lane geometries and intersection layouts. (2) A heuristic-based topology construction module then identifies approach and exit legs and establishes conflict-free lane-to-lane connections. (3) Finally, an automatic simulation-building module converts the extracted network into standard formats, e.g., OpenDRIVE, and generates routable SUMO networks. The framework supports both complete network construction from scratch and local-scale refinement of existing networks through lane-count correction, transition recovery, and geometric regularization. The proposed pipeline provides a practical pathway to generate traffic simulation networks from satellite imagery, significantly reducing manual reconstruction effort and enabling scalable, continuously updated transportation digital twins.

Guo, Hetian [University of Georgia, Athens] (ORCID↗

Examining Rail Transportation Route of Crude Oil in the United States Using Crowdsourced Social Media Data

Safety issues associated with transporting crude oil by rail have been a concern since the boom of the U.S. domestic shale oil production in 2012. During the last decade, over 300 crude-oil-by-rail incidents have occurred in the United States. Some of them have caused adverse consequences including fire and hazardous materials leakage. However, only limited information on crude-on-rail routes and their associated risks is available to the public. To this end, this study proposed an unconventional way to reconstruct crude-on-rail routes using geotagged photos harvested from the Flickr website. The proposed method linked the geotagged photos of crude oil trains posted online with national railway networks to identify potential railway segments that those crude oil trains were traveling on. Here, a shortest path-based method was applied to infer the complete crude-on-rail routes, by utilizing the confirmed railway segments as well as their directional information. Validation of the inferred routes was performed using a public map and official crude oil incident data. The results suggested that the inferred routes based on geotagged photos had high coverage, with approximately 96% of the documented crude oil incidents aligned with the reconstructed crude-on-rail network. The inferred crude oil train routes were found to pass through several metropolitan areas of high population density, who were exposed to potential risk. These findings could improve situational awareness for policy makers and transportation planners. In addition, with the inferred routes, this study has established a good foundation for future crude oil train risk-analyses along the rail route.

42 ENGINEERING↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

Urban Versus Lake Impacts on Heat Stress and Its Disparities in a Shoreline City

Abstract Shoreline cities are influenced by both urban‐scale processes and land‐water interactions, with consequences on heat exposure and its disparities. Heat exposure studies over these cities have focused on air and skin temperature, even though moisture advection from water bodies can also modulate heat stress. Here, using an ensemble of model simulations covering Chicago, we find that Lake Michigan strongly reduces heat exposure (2.75°C reduction in maximum average air temperature in Chicago) and heat stress (maximum average wet bulb globe temperature reduced by 0.86°C) during the day, while urbanization enhances them at night (2.75 and 1.57°C increases in minimum average air and wet bulb globe temperature, respectively). We also demonstrate that urban and lake impacts on temperature (particularly skin temperature), including their extremes, and lake‐to‐land gradients, are stronger than the corresponding impacts on heat stress, partly due to humidity‐related feedback. Likewise, environmental disparities across community areas in Chicago seen for skin temperature are much higher (1.29°C increase for maximum average values per $10,000 higher median income per capita) than disparities in air temperature (0.50°C increase) and wet bulb globe temperature (0.23°C increase). The results call for consistent use of physiologically relevant heat exposure metrics to accurately capture the public health implications of urbanization.

54 ENVIRONMENTAL SCIENCES↗

Linkages Between Mineral Element Composition of Soils and Sediments With Hyporheic Zone Dissolved Organic Matter Chemistry Across the Contiguous United States

The hyporheic zone is a hotspot for biogeochemical cycling where interactions with mineral metals preserve the release and biodegradation of organic matter (OM). A small fraction of OM can still be exchanged between localized sediments and the overlying water column, and recent evidence suggests there exists a longitudinal structuring in sediment dissolved OM (DOM) chemistry across the continental United States (CONUS). In this study, we tested a hypothesis that water extractable sediment DOM chemistry could be explained by sediment metal contents and integrative watershed scale features at the CONUS scale. Crowdsourced samples were characterized for high resolution mass spectrometry and coupled with sediment metals determined via x-ray fluorescence as well as with land cover and soil elemental information obtained from national databases. Our results highlight weak relationships between DOM chemistry and elemental composition at the CONUS scale indicating limited transferability of organo-metal linkages into multi-scale hydrobiogeochemical models.

58 GEOSCIENCES↗

Labeling sequential data from noisy annotations

Crowdsourcing algorithms often work under the assumption that the data samples are independent. Recent work has shown that data dependence, such as temporal correlations in sequential data, can be leveraged to improve the label quality. Existing methods that exploit this special structure rely on third-order statistics of the annotator outputs to ensure the identifiability of key latent parameters, which are costly to acquire. This work proposes an approach for integrating crowdsourced annotations under the Dawid-Skene/Hidden Markov Model (DS-HMM) for sequential data based on second-order statistics, which naturally enjoys a lower sample complexity. An effective algorithm is proposed to tackle the challenging optimization problem associated with the proposed estimator. Numerical experiments showcase the effectiveness of the data labeling paradigm.

Marrinan, Timothy P.↗

Estimating building occupancy: a machine learning system for day, night, and episodic events

Building occupancy research increasingly emphasizes understanding the social and physical dynamics of how people occupy space. Opportunities in the open source domain including social media, Volunteered Geographic Information, crowdsourcing, and sensor data have proliferated, resulting in the exploration of building occupancy dynamics at varying spatiotemporal scales. At Oak Ridge National Laboratory, research into building occupancies through the development of a global learning framework that accommodates exploitation of open source authoritative sources, including governmental census and surveys, journal articles, real estate databases, and more, to report national and subnational building occupancies across the world continues through the Population Density Tables (PDT) project. This probabilistic learning system accommodates expert knowledge, experience, and open-source data to capture local, socioeconomic, and cultural information about human activity. It does so through a systematic process of data harmonization techniques in the development of observation models for over 50 building types to dynamically update baseline estimates and report probabilistic diurnal and episodic building occupancy estimates. This discussion will explore how PDT is implemented at scale and expanded based on the development of observation model classes and will explain how to interpret and spatially apply the reported probability occupancy estimates and uncertainty.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Determining the biogeochemical transformations of organic matter composition in rivers using molecular signatures

Inland waters are hotspots for biogeochemical activity, but the environmental and biological factors that govern the transformation of organic matter (OM) flowing through them are still poorly constrained. Here we evaluate data from a crowdsourced sampling campaign led by the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) consortium to investigate broad continental-scale trends in OM composition compared to localized events that influence biogeochemical transformations. Samples from two different OM compartments, sediments and surface water, were collected from 97 streams throughout the Northern Hemisphere and analyzed to identify differences in biogeochemical processes involved in OM transformations. By using dimensional reduction techniques, we identified that putative biogeochemical transformations and microbial respiration rates vary across sediment and surface water along river continua independent of latitude (18°N–68°N). In contrast, we reveal small- and large-scale patterns in OM composition related to local (sediment vs. water column) and reach (stream order, latitude) characteristics. These patterns lay the foundation to modeling the linkage between ecological processes and biogeochemical signals. We further showed how spatial, physical, and biogeochemical factors influence the reactivity of the two OM pools in local reaches yet find emergent broad-scale patterns between OM concentrations and stream order. OM processing will likely change as hydrologic flow regimes shift and vertical mixing occurs on different spatial and temporal scales. As our planet continues to warm and the timing and magnitude of surface and subsurface flows shift, understanding changes in OM cycling across hydrologic systems is critical, given the unknown broad-scale responses and consequences for riverine OM.

rivers↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Maximum respiration rates in hyporheic zone sediments are primarily constrained by organic carbon concentration and secondarily by organic matter chemistry

Abstract. River corridors are fundamental components of the Earth system, and their biogeochemistry can be heavily influenced by processes in subsurface zones immediately below the riverbed, referred to as the hyporheic zone. Within the hyporheic zone, organic matter (OM) fuels microbial respiration, and OM chemistry heavily influences aerobic and anaerobic biogeochemical processes. The link between OM chemistry and respiration has been hypothesized to be mediated by OM molecular diversity, whereby respiration is predicted to decrease with increasing diversity. Here we test the specific prediction that aerobic respiration rates will decrease with increases in the number of unique organic molecules (i.e., OM molecular richness, as a measure of diversity). We use publicly available data across the United States from crowdsourced samples taken by the Worldwide Hydrobiogeochemical Observation Network for Dynamic River Systems (WHONDRS) consortium. Our continental-scale analyses rejected the hypothesis of a direct limitation of respiration by OM molecular richness. In turn, we found that organic carbon (OC) concentration imposes a primary constraint over hyporheic zone respiration, with additional potential influences of OM richness. We specifically observed respiration rates to decrease nonlinearly with the ratio of OM richness to OC concentration. This relationship took the form of a constraint space with respiration rates in most systems falling below the constraint boundary. A similar, but slightly weaker, constraint boundary was observed when relating respiration rate to the inverse of OC concentration. These results indicate that maximum respiration rates may be governed primarily by OC concentration, with secondary influences from OM richness. Our results also show that other variables often suppress respiration rates below the maximum associated with the richness-to-concentration ratio. An important focus of future research will identify physical (e.g., sediment grain size), chemical (e.g., nutrient concentrations), and/or biological (e.g., microbial biomass) factors that suppress hyporheic zone respiration below the constraint boundaries observed here.

58 GEOSCIENCES↗

Scripts and data associated with a manuscript linking soil and sediment elemental composition with dissolved organic matter chemistry across CONUS

This data package provides scripts and geochemical data for a manuscript titled “Linkages between mineral element composition of soils and sediments with hyporheic zone dissolved organic matter chemistry across the contiguous United States” (preprint: doi: 10.22541/essoar.169447343.31694990/v1). This data is associated with the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS, https://whondrs.pnnl.gov) and is an extension of the Summer 2019 Sampling campaign which crowdsourced samples from rivers and sediment across the continental United States. Data from this study can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719. The main objective of this manuscript was to couple sediment water extractable dissolved organic matter chemistry, defined by ultra-high resolution mass spectrometry, with localized sediment elemental composition and watershed scale soil elemental characteristics. This data package contains one main folder with four subfolders. The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) an R markdown to reproduce manuscript figures and analyses; (5) a pdf of instructions to reproduce NGS interpolations with ArcGIS software; and (6) a python script to reproduce NGS extrapolations with python. The four subfolders contain files required to reproduce NGS extrapolations include (1) ‘CONUS_boundaries’ containing boundary layers (.shp) for the Continental United States; (2) ‘ngs_project’ containing files (.shp) with point level NGS soil elemental data (Grossman et al., 2004); (3) ‘raster_outputs’ containing the interpolated raster output files for various soil elements; and (4) ‘NGS_Chemistry_Final’ contain final extracted soil elemental data.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗