Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

GLORIA - A Globally Representative Hyperspectral In Situ Dataset for Optical Sensing of Water Quality

The development of algorithms for remote sensing of water quality (RSWQ) requires a large amount of in situ data to account for the bio-geo-optical diversity of inland and coastal waters. The GLObal Reflectance community dataset for Imaging and optical sensing of Aquatic environments (GLORIA) includes 7,572 curated hyperspectral remote sensing reflectance measurements at 1 nm intervals within the 350 to 900 nm wavelength range. In addition, at least one co-located water quality measurement of chlorophyll a , total suspended solids, absorption by dissolved substances, and Secchi depth, is provided. The data were contributed by researchers affiliated with 59 institutions worldwide and come from 450 different water bodies, making GLORIA the de-facto state of knowledge of in situ coastal and inland aquatic optical diversity. Each measurement is documented with comprehensive methodological details, allowing users to evaluate fitness-for-purpose, and providing a reference for practitioners planning similar measurements. We provide open and free access to this dataset with the goal of enabling scientific and technological advancement towards operational regional and global RSWQ monitoring.

remote sensing of water quality↗

Automated Collection of Scientific Publications Linked to NASA Earth Science Datasets

NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.

Infometrics↗

VESIcal: A Critical Approach to Volatile Solubility Modelling Using the Open-Source Engine Vesical

Accurate models of H(2)O and CO(2) solubility in silicate melts are vital for understanding volcanic plumbing systems. These models are used to estimate the depths of magma storage regions from melt inclusion volatile contents, investigate the role of volatile exsolution as a driver of volcanic eruptions, and track the degassing path followed by a magma ascending to the surface. However, despite the large increase in the number of experimental constraints over the last two decades, many recent studies still utilize an earlier generation of models which were calibrated on experimental datasets with restricted compositional ranges. This may be because many of the available tools for more recent models require large numbers of input parameters to be hand-typed (e.g., temperature, concentrations of H(2)O, CO(2), and 8–14 oxides), making them difficult to implement on large datasets. Here, we use a new open-source Python3 tool, VESIcal, to critically evaluate the behaviors and sensitivities of different solubility models for a range of melt compositions. Using literature datasets of andesitic-dacitic experimental products and melt inclusions as case studies, we illustrate the importance of evaluating the calibration dataset of each model. Finally, we highlight the limitations of particular data presentation methods, such as isobar diagrams, and provide suggestions for alternatives, and best practices regarding the presentation and archiving of data. This review will aid the selection of the most applicable solubility model for different melt compositions, and identifies areas where additional experimental constraints on volatile solubility are required.

magma↗

Advances in building data management for building performance standards using the SEED platform

Reducing energy consumption and greenhouse gas emissions in the built environment is a critical step in achieving emission goals to mitigate climate change impacts. Local, federal, and international jurisdictions are deploying several methods to reduce energy and emissions such as voluntary and mandatory benchmarking and building performance standards, requiring building owners to reach energy and emission targets. Jurisdictions leveraging benchmarking and building performance standards require knowledge of the buildings covered; which is a large task due to staffing constraints, limited information on building characteristics and tax parcel data, and the need for advanced data management techniques to align datasets. This paper describes an open-source platform's recent advances to create consistent taxonomies, identify erroneous data, enable auditability, and track building performance. The paper concludes with two use cases on how the platform has been used by jurisdictions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Atomistic Simulation of Glasses and Amorphous Materials: Challenges and Opportunities for the Next Decade

Atomistic simulations have become indispensable tools for understanding glass structure, dynamics, and properties, yet persistent challenges limit their predictive power. This perspective examines three interconnected issues, namely glass formation procedures, interatomic potential development, and machine learning applications, which emerged from the 5th International Workshop on Challenges of Atomistic Simulations of Glasses and Amorphous Materials. We identify convergent community priorities for (i) standardized validation protocols, (ii) curated benchmark datasets with complete metadata, and (iii) open repositories for glasses. A systematic was forward is provided by a hierarchical validation framework for assessing the structural fidelity, property prediction, and behavioral realism of simulation techniques. Looking ahead, transformative advances are promised by the fusion of classical techniques with machine learning based approaches, for instance, by integrating swap Monte Carlo with machine-learning (ML) potentials, leveraging foundation models through transfer learning, and finetuning ML potentials with experimental data. Progress depends on the community committing to validated models, reproducible protocols, and sustained data sharing.

Krishnan, N. M. Anoop↗

A Million Person Study Innovation: Evaluating Cognitive Impairment and other Morbidity Outcomes from Chronic Radiation Exposure Through Linkages with the Centers for Medicaid and Medicare Services Assessment and Claims Data

Here, the study of One Million U.S. Radiation Workers and Veterans, the Million Person Study (MPS), examines the health consequences, both cancer and non-cancer, of exposure to ionizing radiation received gradually over time. Recently the MPS has focused on mortality patterns from neurological and behavioral conditions, e.g., Parkinson's disease, Alzheimer's disease, dementia, and motor neuron disease such as amyotrophic lateral sclerosis. A fuller picture of radiation-related late effects comes from studying both mortality and the occurrence (incidence) of conditions not leading to death. Accordingly, the MPS is identifying neurocognitive diagnoses from fee-for-service insurance claims from the Centers for Medicare and Medicaid Services (CMS), among Medicare beneficiaries beginning in 1999 (the earliest date claims data are available). Linkages to date have identified ∼540,000 workers with available health information. Such linkages provide individual information on important co-factor and confounding variables such as smoking, alcohol consumption, blood pressure, obesity, diabetes and many other health and demographic characteristics. The total person-level set of time-dependent variables, outcomes, organ-specific dose measures, co-factors, and demographics will be massive and much too large to be evaluated with standard software. Thus, development of specialized open-source software designed for large datasets (Colossus) is nearly complete. The wealth of information available from CMS claims data, coupled with individual dose reconstructions, will thus greatly enhance the quality and precision of health evaluations for this new field of low-dose radiation and neurocognitive effects.

Dauer, Lawrence T.↗

A reproducible study design for the MIMIC-IV in-hospital mortality task

Open, tabular electronic health record (EHR) datasets such as MIMIC-III and MIMIC-IV have become critical resources for developing machine learning (ML) models addressing clinical prediction tasks, including hospital readmission, length of stay, and in-hospital mortality (IHM). While MIMIC-III has benefited from well-established preprocessing pipelines and standardized feature sets, MIMIC-IV remains comparatively challenging to work with because there are no standardized benchmarks to support reproducibility and comparability across studies. To address this limitation, we present a rigorously curated MIMIC-IV custom feature set optimized for IHM prediction, constructed through a reproducible preprocessing pipeline and feature selection strategy.

97 MATHEMATICS AND COMPUTING↗

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING↗

Temporal Stability of the NDVI-LAI Relationship in a Napa Valley Vineyard

Remotely sensed normalized difference vegetation index (NDVI) values, derived from high-resolution satellite images, were compared with ground measurements of vineyard leaf area index (LAI) periodically during the 2001 growing season. The two variables were strongly related at six ground calibration sites on each of four occasions (r squared = 0.91 to 0.98). Linear regression equations relating the two variables did not significantly differ by observation date, and a single equation accounted for 92 percent of the variance in the combined dataset. Temporal stability of the relationship opens the possibility of transforming NDVI maps to LAI in the absence of repeated ground calibration fieldwork. In order to take advantage of this circumstance, however, steps should be taken to assure temporal consistency in spectral data values comprising the NDVI.

Johnson, L. F.↗

Multi-Mission Terrain Classifier for Safe Rover Navigation and Automated Science

We previously presented Soil Property and Object Classification (SPOC), a machine learning-based terrain classifier for Mars rovers, for automatically segmenting rover images by its surface type such as sand and bedrock. This paper presents a number of practical improvements to pave the way for potential future onboard deployment. First, we achieved 97.0% overall pixel accuracy, evaluated against the classification generated by human experts on images from Mars Science Laboratory (MSL) missions. The substantial increase in accuracy was primarily enabled by the sheer volume of data used for training; we created a new large-scale dataset of Martian terrain labels, namely AI4Mars, which contains more than 400k labels contributed by citizen scientists for 50k images taken by the Mars Exploration Rovers (MER) and Mars Science Laboratory (MSL) rover. Second, we demonstrated that SPOC can quickly adapt to a new mission landed on a previously unseen site. Specifically, we pretrained a model with MER and MSL data from the AI4Mars dataset and then adapted to the Mars 2020 Rover (M2020) by feeding a small volume of data between Sol 0 and 157; the adapted model was tested on Sol 200-203 and resulted in 84.2% overall pixel accuracy and 93.4% reliability (recall) for detecting sand, the most concerning class for rover’s traversability. Third, we found that pretraining can substantially mitigate the decline of accuracy over time. We showed that the performance of a SPOC model pretrained with the ImageNet dataset and then trained by MSL images only up to Sol 390 remains comparable to a model trained by images up to Sol 1689 on the test data after Sol 1689. Fourth, we reimplemented SPOC with a light-weight convolutional neural network (CNN), MobileNetV2, which typically runs within tens of milliseconds (ms) on mobile processors such as Qualcomm’s Snapdragon. Finally, we released the AI4Mars dataset to the public to encourage open innovation.

Ono, Masahiro↗

Satellite Retrieval of Atmospheric Water Budget over Gulf of Mexico- Caribbean Basin: Seasonal Variability

This study presents results from a multi-satellite/multi-sensor retrieval system designed to obtain the atmospheric water budget over the open ocean. A combination of hourly-sampled monthly datasets derived from the GOES-8 5 Imager and the DMSP 7-channel passive microwave radiometer (SSM/I) have been acquired for the Gulf of Mexico-Caribbean Sea basin. Whereas the methodology is being tested over this basin, the retrieval system is designed for portability to any open-ocean region. Algorithm modules using the different datasets to retrieve individual geophysical parameters needed in the water budget equation are designed in a manner that takes advantage of the high temporal resolution of the GOES-8 measurements, as well as the physical relationships inherent to the SSM/I passive microwave signals in conjunction with water vapor, cloud liquid water, and rainfall. The methodology consists of retrieving the precipitation, surface evaporation, and vapor-cloud water storage terms in the atmospheric water balance equation from satellite techniques, with the water vapor advection term being obtained as the residue needed for balance. Thus, we have sought to develop a purely satellite-based method for obtaining the full set of terms in the atmospheric water budget equation without requiring in situ sounding information on the wind profile. The algorithm is partly validated by first cross-checking all the algorithm components through multiple-algorithm retrieval intercomparisons. More fundamental validation is obtained by directly comparing water vapor transports into the targeted basin diagnosed from the satellite algorithm to those obtained observationally from a network of land-based upper air stations that nearly uniformly surround the basin. Total columnar atmospheric water budget results will be presented for an extended annual cycle consisting of the months of October-97, January-98, April-98, July-98, October-98, and January-1999. These results are used to emphasize the changing relationship in E-P, as well as in the varying roles of storage and advection in balancing E-P both on daily and monthly time scales and on localized and basin space scales. Results from the algorithm-to-algorithm intercomparisons will also be presented in the context of sensitivity testing to help understand the intrinsic uncertainties in the water budget terms.

Smith, Eric A.↗

Machine Learning meets Algebraic Combinatorics: A Suite of Benchmark Datasets to Accelerate AI for Mathematics Research

The use of benchmark datasets has become an important engine of progress in machine learning (ML) over the past 15 years. Recently there has been growing interest in utilizing machine learning to drive advances in research-level mathematics. However, off-the-shelf solutions often fail to deliver the types of insights required by mathematicians. This suggests the need for new ML methods specifically designed with mathematics in mind. The question then is: what benchmarks should the community use to evaluate these? On the one hand, toy problems such as learning the multiplicative structure of small finite groups have become popular in the mechanistic interpretability community whose perspective on explainability aligns well with the needs of mathematicians. While toy datasets are a useful benchmark for initial work, they lack the scale, complexity, and sophistication of many of the principal objects of study in modern mathematics. To address this, we introduce a new collection of benchmark datasets, Algebraic Combinatorics Benchmarks (ACBench), representing either classic or open problems in algebraic combinatorics, a subfield of mathematics that studies discrete structures arising from abstract algebra. After describing the datasets, we discuss the challenges involved in constructing “good” mathematics benchmarks, describe baseline model performance, and discuss some of the insights these datasets can provide that may be of interest even to those who are not interested in mathematics research itself.

97 MATHEMATICS AND COMPUTING↗

Use of GOES, SSM/I, TRMM Satellite Measurements Estimating Water Budget Variations in Gulf of Mexico - Caribbean Sea Basins

This study presents results from a multi-satellite/multi-sensor retrieval system designed to obtain the atmospheric water budget over the open ocean. A combination of 3ourly-sampled monthly datasets derived from the GOES-8 5-channel Imager, the TRMM TMI radiometer, and the DMSP 7-channel passive microwave radiometers (SSM/I) have been acquired for the combined Gulf of Mexico-Caribbean Sea basin. Whereas the methodology has been tested over this basin, the retrieval system is designed for portability to any open-ocean region. Algorithm modules using the different datasets to retrieve individual geophysical parameters needed in the water budget equation are designed in a manner that takes advantage of the high temporal resolution of the GOES-8 measurements, as well as the physical relationships inherent to the TRMM and SSM/I passive microwave measurements in conjunction with water vapor, cloud liquid water, and rainfall. The methodology consists of retrieving the precipitation, surface evaporation, and vapor-cloud water storage terms in the atmospheric water balance equation from satellite techniques, with the water vapor advection term being obtained as the residue needed for balance. Thus, the intent is to develop a purely satellite-based method for obtaining the full set of terms in the atmospheric water budget equation without requiring in situ sounding information on the wind profile. The algorithm is validated by cross-checking all the algorithm components through multiple- algorithm retrieval intercomparisons. A further check on the validation is obtained by directly comparing water vapor transports into the targeted basin diagnosed from the satellite algorithms to those obtained observationally from a network of land-based upper air stations that nearly uniformly surround the basin, although it is fair to say that these checks are more effective m identifying problems in estimating vapor transports from a leaky operational radiosonde network than in verifying the transport estimates determined from the satellite algorithm system Total columnar atmospheric water budget results are presented for an extended annual cycle consisting of the months of October-97, January-98, April-98, July-98,October-98, and January 1999. These results are used to emphasize the changing relationship in E-P, as well as in the varying roles of storage and advection in balancing E-P both on daily and monthly time scales and on localized and basin space scales. Results from the algorithm-to-algorithm intercomparisons are also presented in the context of sensitivity testing to help understand the intrinsic uncertainties in evaluating the water budget terms by an all-satellite algorithm approach.

Smith, Eric A.↗

Monthly-Diurnal Water Budget Variability Over Gulf of Mexico-Caribbean Sea Basin from Satellite Observations

This study presents results from a multi-satellite/multi-sensor retrieval system design d to obtain the atmospheric water budget over the open ocean. A combination of hourly-sampled monthly datasets derived from the GOES-8 5-channel Imager, the TRMM TMI radiometer, and the DMSP 7-channel passive microwave radiometers (SSM/I) have been acquired for the combined Gulf of Mexico-Caribbean Sea basin. Whereas the methodology has been tested over this basin, the retrieval system is designed for portability to any open-ocean region. Algorithm modules using the different datasets to retrieve individual geophysical parameters needed in the water budget equation are designed in a manner that takes advantage of the high temporal resolution of the GOES-8 measurements, as well as the physical relationships inherent to the TRMM and SSM/I passive microwave measurements in conjunction with water vapor, cloud liquid water, and rainfall. The methodology consists of retrieving the precipitation, surface evaporation, and vapor-cloud water storage terms in the atmospheric water balance equation from satellite techniques, with the water vapor advection term being obtained as the residue needed for balance. Thus, the intent is to develop a purely satellite-based method for obtaining the full set of terms in the atmospheric water budget equation without requiring in situ sounding information on the wind profile. The algorithm is validated by cross-checking all the algorithm components through multiple-algorithm retrieval intercomparisons. A further check on the validation is obtained by directly comparing water vapor transports into the targeted basin diagnosed from the satellite algorithms to those obtained observationally from a network of land-based upper air stations that nearly uniformly surround the basin, although it is fair to say that these checks are more effective in identifying problems in estimating vapor transports from a "leaky" operational radiosonde network than in verifying the transport estimates determined from the satellite algorithm system. Total columnar atmospheric water budget results are presented for an extended annual cycle consisting of the months of October-97, January-98, April-98, July-98,October-98, and January- 1999. These results are used to emphasize the changing relationship in E-P, as well as in the varying roles of storage and advection in balancing E-P both on daily and monthly time scales and on localized and basin space scales. Results from the algorithm-to-algorithm intercomparisons are also presented in the context of sensitivity testing to help understand the intrinsic uncertainties in evaluating the water budget terms by an all-satellite algorithm approach.

Smith, E. A.↗

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields in the last two decades, in part thanks to an increasing culture of open data sharing and reuse. Due to its capability for identifying complex relationships and patterns, AI/ML methodology is particularly well suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are many key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Even with the positive culture of Open Science and data sharing, inexperienced researchers working quickly without proper checks can produce models that perform poorly outside of the immediate training dataset. Lessons learned from biological AI/ML research indicate that Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Andrew Casaletto↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

Validation of Long-Term Global Aerosol Climatology Project Optical Thickness Retrievals Using AERONET and MODIS Data

A comprehensive set of monthly mean aerosol optical thickness (AOT) data from coastal and island AErosol RObotic NETwork (AERONET) stations is used to evaluate Global Aerosol Climatology Project (GACP) retrievals for the period 1995-2009 during which contemporaneous GACP and AERONET data were available. To put the GACP performance in broader perspective, we also compare AERONET and MODerate resolution Imaging Spectroradiometer (MODIS) Aqua level-2 data for 2003-2009 using the same methodology. We find that a large mismatch in geographic coverage exists between the satellite and ground-based datasets, with very limited AERONET coverage of open-ocean areas. This is especially true of GACP because of the smaller number of AERONET stations at the early stages of the network development. Monthly mean AOTs from the two over-the-ocean satellite datasets are well-correlated with the ground-based values, the correlation coefficients being 0.81-0.85 for GACP and 0.74-0.79 for MODIS. Regression analyses demonstrate that the GACP mean AOTs are approximately 17%-27% lower than the AERONET values on average, while the MODIS mean AOTs are 5%-25% higher. The regression coefficients are highly dependent on the weighting assumptions (e.g., on the measure of aerosol variability) as well as on the set of AERONET stations used for comparison. Comparison of over-the-land and over-the-ocean MODIS monthly mean AOTs in the vicinity of coastal AERONET stations reveals a significant bias. This may indicate that aerosol amounts in coastal locations can differ significantly from those in adjacent open-ocean areas. Furthermore, the color of coastal waters and peculiarities of coastline meteorological conditions may introduce biases in the GACP AOT retrievals. We conclude that the GACP and MODIS over-the-ocean retrieval algorithms show similar ranges of discrepancy when compared to available coastal and island AERONET stations. The factors mentioned above may limit the performance of the validation procedure and cause us to caution against a direct extrapolation of the presented validation results to the entirety of the GACP dataset.

MODIS (radiometry)↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS↗