Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

GraphTango: A Hybrid Representation Format for Efficient Streaming Graph Updates and Analysis

Abstract Streaming graph processing performs batched updates and analytics on a time-evolving graph. The underlying representation format of the graph largely determines the throughputs of these updates and analytics phases. Existing representation formats usually employ variations of hash tables or adjacency lists. However, a recent study showed that the adjacency-list-based approaches perform poorly on heavy-tailed graphs, and the hash table-based approaches suffer on short-tailed graphs. We propose GraphTango, a hybrid representation format that provides excellent update and analytics throughput regardless of the graph’s degree distribution. GraphTango dynamically switches among three different formats based on a vertex’s degree: (i) Low-degree vertices store the edges directly with the neighborhood metadata, confining accesses to a single cache line, (2) Medium-degree vertices use adjacency lists, and (3) High-degree vertices use hash tables as well as adjacency lists. In this case, the adjacency list provides fast traversal during the analytics phase, while the hash table provides constant-time lookups during the update phase. We further optimized the performance by designing an open-addressing-based hash table that fully utilizes every fetched cache line. In addition, we developed a thread-local lock-free memory pool that allows fast growing/shrinking of the adjacency lists and hash tables in a multi-threaded environment. We evaluated GraphTango with the help of the SAGA-Bench framework and compared it with four other representation formats: Stinger, Degree-aware Robin Hood Hashing, and two adjacency list-based formats with different workload balancing scheme. On average, GraphTango provides 4.5x higher insertion throughput, 3.2x higher deletion throughput, and 1.1x higher analytics throughput over the next best format. Furthermore, we integrated GraphTango with the state-of-the-art graph processing frameworks DZiG and RisGraph. Compared to the vanilla DZiG and vanilla RisGraph , [ GraphTango + DZiG ] and [ GraphTango + RisGraph ] reduces the average batch processing time by 2.3x and 1.5x, respectively.

Ahmed, Alif↗

Timeseries Photos of a Variably Inundated Stream: Umtanum Creek, Washington, United States

This dataset is associated with a broader study using game camera timeseries photos collected to evaluate stream variable inundation via changes in width (i.e. wet fraction). Four game cameras were deployed along Umtanum Creek (Washington, United States) to track changes in stream inundation over time. Drone imagery was collected at the same location on October 18, 2024 which was used to construct a digital elevation model (DEM) of the streambed topography. The associated paper and data can be found at https://doi.org/10.1016/j.envsoft.2025.106715 (Bao et al., 2025a)) and https://doi.org/10.15485/2589885 (Bao et al., 2025b), respectively. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) files that describes each file and a data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) field protocol; and (5) folders containing game camera photos. Game camera photos are organized into folders for each camera (CDL, CUL, CDR, CUR; see readme for information on camera naming) by the month photos were collected. All files are .csv, .jpg, or .pdf.

AI image segmentation↗

EV Profile Capture

NextGen Profiles' EV profile capture efforts aimed to explore the variance in performance and evaluate how different operational conditions influence production EV charging behavior. Data were collected at a frequency of 10 Hz from both the EV and EVSE during each charge session. These charge session parameters were then entered into a time-series database for further analysis. The data were gathered under different operational conditions to examine the effects of various factors such as battery state of charge, battery temperature, vehicle condition, smart charge management, and EVSE limitations. The EV profile capture dataset includes extensive high-power charging data from 16 different EVs—comprising light-, medium-, and heavy-duty vehicles—along with EVSE from various suppliers. To protect confidentiality, the EV and EVSE metadata are anonymized, and the publicly released datasets are aggregated to 0.1-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Metabolomics (PB-DP5)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Culture samples were collected at 0, 0.5, 1, 2, 4, 6, and 8 hours for extracellular sucrose analysis. Circadian metabolomics data was acquired using a Agilent single quadrupole gas chromatography-mass spectrometer and processed using Agilent Mass Hunter for targeted sucrose quantification. Metabolomic analysis of PCC 7942 light-dark cycle cultures transitioned to constant light revealed distinct temporal patterns in sucrose production. Processed metabolomic datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed GC-MS results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Upland tidal brackish marsh specific conductivity and salinity measurements, PIE LTER, Plum Island Sound, MA, 2023

This dataset includes raw and corrected specific conductivity, temperature, and calculated salinity measurements collected at 10 cm depth in a Typha angustifolia-dominated tidal brackish wetland at the upper estuary of Plum Island Sound in Newbury, Massachusetts (MA) within the Plum Island Ecosystems Long Term Ecological Research (PIE LTER) site. Measurements were taken to evaluate temporal variation in porewater salinity (a proxy for porewater sulfate concentration) in high frequency to assess soil and plant responses to changes in salinity. Measurements were collected using an Onset HOBO U24-002 Saltwater Conductivity/Salinity data logger deployed in a well. Specific conductivity was corrected using non-linear temperature compensation, and salinity was calculated using the Practical Salinity Scale 1978 via Onset's HOBOware software. Reference conductivity measurements to correct for sensor drift were taken at the start and end of each deployment using a HACH HQ14D Portable Conductivity Meter. Data were then filtered in MATLAB to remove values logged while the sensor was out of the well or during post-deployment equilibration. Detailed metadata, including variable descriptions, sampling methods, QA/QC procedures, and site information, are provided in the files: Typha_ctd_salinity_dd.csv and Typha_MI_ctd_salinity_2023_flmd.csv.

54 ENVIRONMENTAL SCIENCES↗

Upland tidal brackish marsh specific conductivity and salinity measurements, PIE LTER, Plum Island Sound, MA, May-December 2022

This dataset includes raw and corrected specific conductivity, temperature, and calculated salinity measurements collected at 10 cm depth in a Typha angustifolia-dominated tidal brackish wetland at the upper estuary of Plum Island Sound in Newbury, Massachusetts (MA) within the Plum Island Ecosystems Long Term Ecological Research (PIE LTER) site. Measurements were taken to evaluate temporal variation in porewater salinity (a proxy for porewater sulfate concentration) in high frequency to assess soil and plant responses to changes in salinity. Measurements were collected using an Onset HOBO U24-002 Saltwater Conductivity/Salinity data logger deployed in a well. Specific conductivity was corrected using non-linear temperature compensation, and salinity was calculated using the Practical Salinity Scale 1978 via Onset's HOBOware software. Reference conductivity measurements to correct for sensor drift were taken at the start and end of each deployment using a HACH HQ14D Portable Conductivity Meter. Data were then filtered in MATLAB to remove values logged while the sensor was out of the well or during post-deployment equilibration. Detailed metadata, including variable descriptions, sampling methods, QA/QC procedures, and site information, are provided in the files: Typha_ctd_salinity_dd.csv and Typha_MI_ctd_salinity_2023_flmd.csv.

54 ENVIRONMENTAL SCIENCES↗

Specific conductivity and salinity of the Parker River, PIE LTER, Plum Island Sound MA, August-November 2022

This dataset contains specific conductivity and calculated salinity data of Parker River water at a tidal brackish wetland dominated by Typha angustifolia at the upper estuary of the Plum Island Sound in Newbury, Massachusetts (MA) within the Plum Island Ecosystems Long Term Ecological Research site (PIE LTER). Measurements were taken to evaluate temporal changes in surface water salinity in high frequency to characterize boundary conditions of soil and plant responses to changes in salinity. A PVC pipe was installed in a low elevation spot in the creek bank so that the bottom of the pipe sat on the sediment surface allowing flushing with water during flooding. Raw measurements were collected using an Onset HOBO U24-002 Saltwater Conductivity/Salinity data logger. The specific conductance and salinity measurements were corrected and calculated respectively using Onset’s HOBOware software and reference specific conductivity measurements taken in tandem with the first and last points recorded by the HOBO sensor. These reference measurements were taken using a HACH HQ14D Portable Conductivity Meter. Because of the installation design, only data one hour before and after high tide are used. Metadata files Typha_ctd_salinity_dd.csv and Typha_ctd_salinity_flmd.csv contain detailed information on data variables, sampling and QA/QC methods, and site location.

54 ENVIRONMENTAL SCIENCES↗

Specific conductivity and salinity of the Parker River, PIE LTER, Plum Island Sound MA, March-November 2023

This dataset contains specific conductivity and calculated salinity data of Parker River water at a tidal brackish wetland dominated by Typha angustifolia at the upper estuary of the Plum Island Sound in Newbury, Massachusetts (MA) within the Plum Island Ecosystems Long Term Ecological Research site (PIE LTER). Measurements were taken to evaluate temporal changes in surface water salinity in high frequency to characterize boundary conditions of soil and plant responses to changes in salinity. A PVC pipe was installed in a low elevation spot in the creek bank so that the bottom of the pipe sat on the sediment surface allowing flushing with water during flooding. Raw measurements were collected using an Onset HOBO U24-002 Saltwater Conductivity/Salinity data logger. The specific conductance and salinity measurements were corrected and calculated respectively using Onset’s HOBOware software and reference specific conductivity measurements taken in tandem with the first and last points recorded by the HOBO sensor. These reference measurements were taken using a HACH HQ14D Portable Conductivity Meter. Because of the installation design, only data one hour before and after high tide are used. Metadata files Typha_ctd_salinity_dd.csv and Typha_ctd_salinity_flmd.csv contain detailed information on data variables, sampling and QA/QC methods, and site location.

54 ENVIRONMENTAL SCIENCES↗

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator↗

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

CO2 and CH4 leaf-level fluxes and soil porewater concentrations from common vegetation patches in Louisiana’s coastal wetlands

This dataset contains leaf-level flux and soil porewater concentration measurements of carbon dioxide (CO2) and methane (CH4 ) in plots in the footprint of Ameriflux sites US-LA2 and US-LA3. Leaf fluxes in US-LA2 were measured on patches dominated by Sagittaria lancifolia and co-dominated by Sagittaria lancifolia and Typha latifolia. In US-LA3, fluxes were measured from distinct Juncus roemerianus and Spartina alterniflora patches. The porewater concentrations were collected across a vertical profile (~50 cm depth) at centric locations within 25 m2 plots where we measured the leaf fluxes. US-LA3 included an additional set of measurements in open water spots. We aimed to evaluate differences in leaf fluxes and porewater pools of CO2 and CH4 of representative ecohydrological patches across a salinity gradient. We also used this dataset to help develop ELM-Wet, a more realistic representation of wetland carbon biogeochemical processes within the U.S. Department of Energy’s Energy Exascale Earth System Model (E3SM) Land Model version 1 (ELM v.1). The files can be opened with regular text editors or spreadsheet programs. Version 2.0 (8/26/2025): This is the latest version of this dataset. The update includes additional samples of soil porewater CH4/CO2 concentrations from June-2021 to November-2022, as well as minor adjustments made to V1 samples via changing Henry's solubility to account for porewater salinity. Additionally leaf-level measurments of spectral indices, PSRI, NDVI, and PRI have been added to complement Leaf-level flux measurements. All V1 data sets have been integrated into V2 sheets, ensuring data from the previous version is contained with the additional samples and consistent with V2 metadata.

54 ENVIRONMENTAL SCIENCES↗

Monitoring of ground water table depth and soil moisture at the Point Reyes field site

Ground water table (GWT) depth and soil moisture (SM) have been monitored at several locations at the Point Reyes field site (Californian coastal grassland) from 2021 to 2024. Monitoring is still on-going and data may be added to this archive at later time. The SM data have been acquired using Teros 12 Meter soil moisture sensors placed at 10, 30, 60 and 90 cm depth at 5 locations along a small hillslope. These sensors also collect soil temperature and bulk conductance. In addition, some collocated sensors provide pore pressure and Photochemical Reflectance Index (PRI). The GWT depth has been inferred from various type of Onset pressure transducers. The pressure measurements have been corrected for atmospheric pressure variations and sensor position relative to the ground surface to infer GWT depth, as well as with RTK GPS data to infer GWT elevation. The GWT data have been acquired at 5 distinct locations from 2020 to 2024 with the sensors placed at about 4 m depth. In addition, GWT data has been acquired for the 2023-2024 period with sensors located in 1 m deep shallow wells installed near each deeper well. This data is intended to evaluate possibly different dynamic in shallow (perched) and deep aquifer. The datasets are all provided in csv format. Please note that the interpretation of the GWT data needs to be done with consideration of environmental and well characteristics at the site and uncertainty in various variables. For more information on GWT and SM data, please contact the author.

54 ENVIRONMENTAL SCIENCES↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

CHESS 2025: Discrete-return LiDAR point clouds from NEON AOP surveys

This dataset provides Level 1 (L1) discrete-return light detection and ranging (LiDAR) point cloud data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary unclassified discrete-return LiDAR data delivered by NEON and are provided per flightline as LASzip (LAZ) 1.4 Format 6 files. Data were processed following the workflow described in the NEON L0-to-L1 Discrete Return LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022). Each record in the unclassified point clouds represents a geolocated laser target/return recorded by the LiDAR system, with values for X, Y, Z position and return intensity. All point coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Flight metadata describing flightline boundaries and positional uncertainty by point are also included. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗