Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata validation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

53 records · Page 3

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

Common practices for quantifying methane emissions from plumes detected by remote sensing

This document provides a set of community-accepted practices for quantifying methane emissions based on plumes detected via spectroscopic remote sensing. Its primary goal is to promote consistency in the generation, validation, reporting, and quality assessment of methane emission estimates derived from remote sensing radiances. Developed by subject matter experts with deep experience across all stages of the measurement process, this guidance reflects a critical evaluation of current methodologies and highlights key practices needed to produce reliable, interoperable, and traceable products. The focus is specifically on methane emissions quantified from distinct plumes originating from localized sources, rather than diffuse emissions spread over large regions, which are beyond the scope of this work. This document is intended to serve both data producers and users. For producers, it offers a framework for aligning with field-recognized standards to ensure their outputs meet rigorous quality and transparency criteria. For users, it provides a reference to assess dataset fitness-for-purpose by highlighting essential metadata, assumptions, and methodological choices that underpin emission estimates. By fostering a shared understanding of best practices, this work aims to enhance comparability, confidence, and utility of remotely sensed methane emission products.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” (v3)

This data package is associated with the publication “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” submitted to Journal of Advances in Modeling Earth Systems (Butler et al. 2025). This study developed the Sequential Precipitation Input Tagging (SPIT) framework to tag input precipitation and estimate water transit times and hydrologic tracers. SPIT tags all precipitation events at regular intervals over an extended period (monthly tags over seven years) in a hydrologic model from 2016-2022. SPIT is applied at six National Ecological Observatory Network (NEON) sites across the continental United States to calculate transit time distributions (TTD) and derive from these mean transit times (MTT), fractions of young water (Fyw), and hydrologic tracer concentrations in stream water (δ18O) within a water-tagging enabled version of the Weather Research and Forecast (WT-WRF-Hydro) model with national water model (NWM) configurations. We go on to validate WT-WRF-Hydro estimates against Butler et al. (2023), who analyzed the same NEON sites using stable water isotope data to estimate water transit times. This new tracking method provides a detailed picture of water movement and helps improve predictions about water availability in the future. This data package was originally published in January 2025. It was updated May 2025 (v2; new and modified files) and October 2025 (v3; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. This data package contains the data and scripts used to develop the SPIT framework WT-WRF-Hydro (Water Tagging Weather Research and Forecasting Hydrologic) model and is associated with the following GitHub repository: https://github.com/zbutler33/SPIT-Framework. This data package contains five parent folders: (1) “Manipulated_outputs”, (2) “Metadata”, (3) “Observed”, (4) “Outputs”, and (5) “Scripts”. Each of these parent folders contains additional subfolders and files. Please see the FLMD (“v*_Butler_2024_WT_WRF_Hydro_flmd.csv”) for a list of all the files contained in this data package and descriptions for each. See the data dictionary (“v*_Butler_2024_WT_WRF_Hydro_dd.csv”) for definitions and units of all of the tabular (files ending in “.csv” and ".tsv") column headers.

54 ENVIRONMENTAL SCIENCES↗

Development of a Digital Twin for Hydrogen Dispersion and Safety Assessment in an Electrolyzer Based Hydrogen Production Facility

Digital twin models are virtual representations of physical systems that use real-time data to simulate and optimize performance. This study presents the development and initial implementation of a digital twin (DT) for the electrolyzer-based hydrogen production facility at NREL's Advanced Research on Integrated Energy Systems (ARIES), focused on enhancing safety and optimizing sensor placement through physics-based simulations and metadata integration. The DT incorporates detailed facility-specific information, including component layout, leak locations, and controlled release parameters, to model hydrogen dispersion under varying environmental conditions. Using steady-state computational fluid dynamics (CFD) simulations informed by real meteorological data, such as wind speed, direction, and vertical wind profiles, the DT enables visualization of hydrogen plume behavior and spatial concentration distributions. Comparative analysis between high and low wind speed scenarios illustrates the significant influence of wind dynamics on plume shape and extent, with horizontal momentum dominating dispersion at higher speeds, while buoyancy effects become more prominent under low wind conditions. These simulations generate a rich dataset embedded within the DT, allowing users to assess potential leak outcomes and identify optimal sensor locations based on concentration thresholds. The model supports scenario-based analysis to guide safety strategies and equipment deployment for open-area hydrogen infrastructure. The digital twin thus serves as a dynamic platform for virtual prototyping, providing predictive insight into hydrogen behavior and enhancing risk-informed decision-making. This initial phase establishes a validated foundation for future integration of transient, uncontrolled leak scenarios and real-time sensor feedback, positioning the DT as a critical tool for safety design, operational planning, and adaptive monitoring in hydrogen systems. Overall, the approach demonstrates the value of combining environmental data with digital simulations to inform safer and more efficient deployment of hydrogen technologies.

08 HYDROGEN↗

Development of a Digital Twin for Hydrogen Dispersion and Safety Assessment in an Electrolyzer-Based Hydrogen Production Facility: Preprint

Digital twin models are virtual representations of physical systems that use real-time data to simulate and optimize performance. This study presents the development and initial implementation of a digital twin (DT) for the electrolyzer-based hydrogen production facility at the National Renewable Energy Laboratory (NREL)'s Advanced Research on Integrated Energy Systems (ARIES), focused on enhancing safety and optimizing sensor placement through physics-based simulations and metadata integration. The DT incorporates detailed facility-specific information, including component layout, leak locations, and controlled release parameters, to model hydrogen dispersion under varying environmental conditions. Using steady-state computational fluid dynamics (CFD) simulations informed by real meteorological data, such as wind speed, direction, and vertical wind profiles, the DT enables visualization of hydrogen plume behavior and spatial concentration distributions. Comparative analysis between high and low wind speed scenarios illustrates the significant influence of wind dynamics on plume shape and extent, with horizontal momentum dominating dispersion at higher speeds, while buoyancy effects become more prominent under low wind conditions. These simulations generate a rich dataset embedded within the DT, allowing users to assess potential leak outcomes and identify optimal sensor locations based on concentration thresholds. The model supports scenario-based analysis to guide safety strategies and equipment deployment for open-area hydrogen infrastructure. The digital twin thus serves as a dynamic platform for virtual prototyping, providing predictive insight into hydrogen behavior and enhancing risk-informed decision-making. This initial phase establishes a validated foundation for future integration of transient, uncontrolled leak scenarios and real-time sensor feedback, positioning the DT as a critical tool for safety design, operational planning, and adaptive monitoring in hydrogen systems. Overall, the approach demonstrates the value of combining environmental data with digital simulations to inform safer and more efficient deployment of hydrogen technologies.

08 HYDROGEN↗

Legacy Survey of Space and Time Data Preview 2: visit_table dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the visit_table dataset type. These are metadata, including dates and filters for every visit. This release contains 1 dataset of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 2: visit_summary dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF- DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single- visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the visit_summary dataset type. These are metadata summarizing a visit. This release contains 28,698 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 2: Visit searchable catalog

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of a searchable catalog named Visit. This catalog contains metadata, including dates and filters for every visit. This catalog contains 28,698 rows with 15 columns.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 2: visit_detector_table dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the visit_detector_table dataset type. These are per-detector visit metadata. This release contains 1 dataset of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 2: VisitDetector searchable catalog

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of a searchable catalog named VisitDetector. This catalog contains per-detector visit metadata. This catalog contains 5,150,198 rows with 56 columns.

79 ASTRONOMY AND ASTROPHYSICS↗

Daily evapotranspiration changes during heatwaves at 32 NEON sites, 2019-2021

This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.

Day of Heatwave↗

High-resolution leaf area index maps generated from unoccupied aerial system, Teller Mile 27, Seward Peninsula, Alaska

Leaf area index (LAI), a measure of the amount of one-side leaf area per ground unit, is an important indicator of plant carbon, energy, and water cycle. In the heterogeneous Arctic landscapes, it has been challenging to accurately measure LAI across species and space needed for Earth system model validation. Here, we use multispectral unoccupied aerial systems (UASs) to scale up and map leaf area index (LAI) , in a low-Arctic tundra landscape on the Seward Peninsula, Alaska. We linked previous published LAI measurements with high-resolution, UAS-collected multispectral data collected over the region of Next Generation Ecosystem Experiments in the Arctic (NGEE Arctic)’s Teller Mile Maker 27 site in 2022 to develop random forest (RF) machine learning models to predict and map LAI. 100 RF models were developed to account for uncertainties in ground LAI plot measurements and process scaling. This dataset includes a raster (*.tif) map of the mean LAI value of the 100 RF models, a raster (*.tif) map of the standard deviation of the RF-modeled LAI data, and a user guide (*.pdf).

54 ENVIRONMENTAL SCIENCES↗

Site and endmember spectra of terrestrial vegetation and soils for the Colorado Headwaters Ecological Spectroscopy Study, June-July 2025

This dataset provides site and endmember spectra collected during the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign. The site spectra were collected to help validate airborne hyperspectral data acquired by the National Ecological Observatory Network's aerial observation platform (NEON AOP). Endmember spectra were collected to augment existing spectral libraries with additional samples of bare surfaces and non-photosynthetic vegetation. All measurements were acquired with an Analytical Spectral Devices (ASD) FieldSpec4 Hi-Res NG (Next Generation) spectroradiometer, which records radiance at 1nm (nanometer) intervals from the ultraviolet to the short-wave infrared (350-2500 nm). The dataset includes spectra measured at meadow sites where the CHESS team also collected vegetation samples for trait analyses. The site spectra were collected with the ASD FieldSpec4 palm grip attachment using an 8° field-of-view foreoptic. Site spectra are integrated measurements of the entire surface within the foreoptic’s field of view. For site-level spectra, the sun is the illumination source. A Spectralon panel mounted on a tripod was used for instrument optimization and white reference measurements for all site spectra. Site spectra were acquired within two hours of solar noon and within 48 hours of a NEON AOP overflight. Site spectra are labeled by date, sampling area, and site number according to the naming conventions of the CHESS campaign’s data management plan. The dataset also contains endmember spectra in the following categories: photosynthetic vegetation (PV), non-photosynthetic vegetation (NPV), bare (soil/rock), and flowers. Endmember measurements were acquired using either the contact probe or the leaf clip attachments of the ASD FieldSpec4. In these configurations, the bulb inside the spectrometer provides the light source for the measurements. The spectrometer was optimized and white reference measurements were recorded using the circular white pucks attached to the contact probe and leaf clip. Because they do not rely on solar illumination, contact probe and leaf clip measurements were collected during a broader time frame than the palm grip site spectra. Some endmembers were measured at CHESS meadow sites, while others were collected within the larger sampling area or in nearby locations (e.g. Gothic Townsite) with similar characteristics. Radiance, reflectance, and metadata files are split into three subfolders according to measurement type: proximal/palm grip (prx), contact probe (cp), and leaf clip (lc). Radiance spectra are provided in ASD file format (.asd file extension). All ASD files can be opened using the provided scripts. Metadata is provided in two formats: CSV file format (no geolocation) and GEOJSON file format (includes geolocation for each spectra). The dataset includes a set of pre-processed reflectance spectra as CSV files (yyyymmdd_rfl.csv). The python scripts and jupyter notebook used to calculate reflectance spectra from the ASD radiance data is included here and was previously published at: https://doi.org/10.3334/ORNLDAAC/2446. There is also a folder of JPEG photographs corresponding to selected spectra. We include a protocol document with detailed steps for ASD FieldSpec4 assembly and operations. This data additionally contains a file level metadata (flmd.csv) and data dictionary (dd.csv) file. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

Quality-Controlled Meteorological Data from the Flood Control District of Maricopa County (FCDMC) Network, Phoenix, Arizona (1987-2024)

This dataset contains 15- or 30-minute interval meteorological data from the Flood Control District of Maricopa County (FCDMC), Arizona, USA, covering eight key variables across multiple sensor stations between 1987 and 2024. Each variable is stored as a separate CSV file, containing time-series data that have undergone rigorous quality control (QC) procedures and, where appropriate, short-gap interpolation for consistency. The quality control (QC) pipeline consisted of four sequential tests: (1) a range test to ensure all values fall within physically realistic limits, (2) a step test to identify abrupt and implausible changes between consecutive records, (3) a proximity test that validates flagged values from step test using data from nearby stations and exceedance probability thresholds, and (4) a persistence test to detect and remove periods of unrealistically constant readings. These thresholds were calibrated to Arizona’s environmental conditions and sensor specifications. After QC, short gaps (≤2 hours) were linearly interpolated to ensure consistent temporal resolution, except for wind variables. Due to a major upgrade in FCDMC’s data transmission system, only ALERT-2 protocol data (2016–2024) for wind variables are included; earlier ALERT-1 data were excluded because of irregular sampling and high missing rates. This dataset supports regional climate and infrastructure resilience studies by providing standardized, high-resolution meteorological data for the greater Phoenix metropolitan area.

54 ENVIRONMENTAL SCIENCES↗

Data-model files associated with the manuscript "Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California"

This package contains the data, simulation setups, notebooks and figures used in “Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California” (Xu et al., 2025). In this study, we selected Elkhorn Slough, a tidal estuary, in California, to investigate the impact of wetland restoration and sea level rise on coastal hydrology using the process-based coastal hydrologic model, Advanced Terrestrial Simulator (ATS), informed by site-specific data. We designed a novel modeling workflow for incorporating wetland restoration features into land cover and soil properties for the model parameterization. The validation results demonstrate a strong agreement between modeled and observed data. We studied the characteristics of coastal watershed hydrology, then focused on the surface water dynamics at two wetland sites within Elkhorn Slough, a reference site and a restored site. Our simulation results indicate that the restored site successfully maintains surface elevation, resulting in reduced surface inundation. We also examined the impact of wetland restoration under expected sea level rise over the next few decades. The low-lying Yampah Marsh, the reference site, is likely to be inundated due to future sea level rise when highest tides arrive; while a higher percentage of Hester Marsh, the restored site, would retain marsh vegetation in coming decades, regardless of tidal conditions. Our study provides important information for examining the outcome of restoration practices that include surface elevation in tidal wetlands under climate changes.Several files can be found from this data package.1. README.md: This file describes the title, journal, co-authors, abstract, repository structure and model version.2. Simulation_Setups.zip: The file contains the model configuration files (XML format) for ATS. 3. Notebooks.zip: The file contains the Jupyter notebooks for generating the pre- and post-restoration meshes and the meshes of future scenarios. 4. Figures.zip: The file contains the figures used in the manuscript.5. Data.zip: The file contains the data used to drive the model simulations, including watershed and wetlands boundaries, mesh files and references to additional datasets (e.g., meteorological forcing, tidal dataset, DEMs, land cover, soil properties). Also, it contains water level observations at the restored wetland.

54 ENVIRONMENTAL SCIENCES↗

Effects of 9.5 Years of Whole-Soil Warming on the Fatty Acid and n-Alkanes Composition in Bulk Soil and Density Fractions at Blodgett Experimental Forest, California, USA

Original data of molecular data (fatty acids and n-alkanes) including concentrations and calculated molecular proxies in a whole-soil warming experiment at the Blodgett Forest Research Station after 9.5 years of warming. The study site has a Mediterranean climate with annual average temperature of 12.5 ℃ and annual average precipitation of 1774 mm. The study site is characterized by a mesic Ultic Alfisol formed from granitic parent material, corresponding to a Dystric Cambisol under the World Reference Base for Soil Resources (WRB) classification system. Experimental warming is applied throughout the soil profile to a depth of 1 m using vertically embedded heating cables that raise soil temperature by 4 °C relative to ambient conditions. Soil samples were collected on 1 May 2023, after the experiment had been operating continuously for about 9.5 years since its initiation in January 2014.The data has been processed from raw data and cross-validated by other peers. The dataset includes: - Bulk_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including Carbon Preference Index (CPI) and Average Chain Length (ACL) of bulk soil organic carbon; - Fractions_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including CPI and ACL of free particulate organic matter (fPOM) and mineral-associated organic matter (MAOM); - Bulk_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of bulk soil organic carbon; - Fractions_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of fPOM and MAOM; - n-Alkanes_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the n-alkane monomers identified and integrated for bulk soil, fPOM and MAOM; - Fattyacid_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the fatty acid monomers including diacids identified and integrated for bulk soil, fPOM, and MAOM. All data are provided in CSV format and can be viewed using Microsoft Excel. We specifically look at fatty acids (FA) and n-alkanes in bulk soil, fPOM and MAOM and calculated molecular proxies such as CPI and ACL to understand the source of oragnic carbon (with ACL) and degree of decomposition (CPI) of each soil fraction. Due to lack of long-chain fatty acids (carbon number ⩾ 20), microorganism-derived organic carbon is characterized by shorter ACL in comparison to plant-derived organic carbon. Fresh SOC is characterized by even-over-odd dominance for fatty acids and odd-over-even dominance for n-alkanes. Therefore, CPI indicates whether soil organic carbon (SOC) represents fresh input (CPI > 10) or is strongly decomposed (close to 1). The research questions should be then, after 9.5-year warming: 1. whether the relative contribution between microorganism-derived and plant-derived SOC in each soil fraction? 2. whether fPOM became more decomposed whereas MAOM remained relatively persistent in each soil fraction across the soil depth?

Carbon↗