Search NASA⌕ Search

SEARCH · Search NASA

Results for “Jupyter notebook”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”

This page contains the datasets and code #O5097 Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”. Data literature-derived cultivation experiments for polyhydroxybutyrate (PHB) production in Synechocystis sp. PCC 6803 and were used for ML model development, interpretation, and experimental validation.

Lalonde, Jessica N. [Los Alamos National Laborator↗

Jupyter notebooks for analyzing transmission SAXS/WAXS from beamline 7.3.3 during operando membrane fouling experiments v1.0

This software consists of python-based Jupyter Notebooks for processing transmission x-ray scattering images collected at beamline 7.3.3 at the Advanced Light Source. The measurements considered in these analyses are collected during membrane fouling experiments, where contaminants in the water deposit on/attach to the membrane surface. This software is used to elucidate the mechanisms of membrane fouling that occur during these operando membrane fouling experiments, but the analyses provided in these scripts can be extended to other scientific cases.

Landsman, Matthew [Lawrence Berkeley National Labo↗

A Jupyter Notebook Environment For Multibody Dynamics

DARTS is a rigid/flexible multibody dynamics toolkit for themodeling and simulation of aerospace and robotic vehicles forengineering applications. In this paper we describe an on-line,browser-based environment using Jupyter notebooks to supporttraining needs for the DARTS software. The suite of curated tutorial notebooks is organized into different topic areas, and intomultiple themes within each topic area. The notebooks within atheme use a progression of examples for users to expand theirunderstanding of the software. The topic areas include one onthe DARTS multibody dynamics software and another one on thetheory underlying the multibody dynamics formulation. We alsodescribe a number of Jupyter extensions that were used - andsome developed in house - to enhance the notebook interface foruse with the dynamics simulation software. One significant extension we implemented allows the embedding of live 3D visualizations within simulation notebooks.

Gaut, Aaron↗

Hands-On, Heads-Up: Blending Cyber T&E with Data Science-Driven Training in Jupyter Notebooks

In an era of increasingly sophisticated threats to critical infrastructure, cybersecurity professionals must be more than just aware; they must be immersed, agile, and equipped to operate in environments where failure is not an option. Nowhere is this truer than in the nuclear sector, where cyber-physical systems, regulatory scrutiny, and insider threat potential demand a new generation of hands-on, technically fluent defenders. This paper presents a unified training approach that integrates Cybersecurity Test and Evaluation (T&E) with data science techniques using Jupyter Notebooks as the interactive lab environment. The program centers on a modular, scenario-driven curriculum designed to build not just knowledge but practical capability in the assessment and defense of radiation detection systems, firmware interfaces, and operational security postures.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗

SAGE III/ISS Rapid Data Analysis Through Dashboarding with Jupyter Notebooks

Spaceborne remote sensing observations of Earth’s atmosphere produce significant quantities of data over the life of each mission. In the case of the Stratospheric Aerosol and Gas Experiment III on the International Space Station (SAGE III/ISS) nearly four years of vertical profiles of atmospheric ozone, water vapor, and nitrogen dioxide concentrations as well as aerosol extinction coefficients have been released. The dichotomy of the desire for both long-term trends in the atmospheric state alongside the assessment of short-term impacts of major disruptive events such as volcanic eruptions and pyrocumulus injections requires agile tools to handle these cases in near real-time as new data are produced. The analysis landscape is further complicated by the desire to compare results between the numerous contemporary observations available for a given dataset. The SAGE III/ISS team has developed a suite of tools leveraging modern web-based frameworks allowing members to interact with a dashboard-style interface to load the data record, assess new profiles as they are generated and in ensemble, compare between species, and additionally add in measurements observed by other platforms as necessary. Leveraging a commonly packaged data format of NetCDF alongside the Python Jupyter Notebook framework, the data can be served to interested parties from an analysis server while still runnable on personal systems if required. This presentation illustrates the ecosystem developed by the SAGE III/ISS team, the applicability to measurements made by any limb-observing platform, and the benefit to transforming routine analyses into readily accessible dynamic plots. Frameworks currently exist at larger scales with projects such as GIOVANNI, and this illustration seeks to show that similar frameworks are accessible and possible within the local research environment while simultaneously unloading human processing cycles for more specialized analysis tasks.

Dashboarding↗

GES DISC Data Recipes in Jupyter Notebooks

The Earth Science Data and Information System (ESDIS) Project manages twelve Distributed Active Archive Centers (DAACs) which are geographically dispersed across the United States. The DAACs are responsible for ingesting, processing, archiving, and distributing Earth science data produced from various sources (satellites, aircraft, field measurements, etc.). In response to projections of an exponential increase in data production, there has been a recent effort to prototype various DAAC activities in the cloud computing environment. This, in turn, led to the creation of an initiative, called the Cloud Analysis Toolkit to Enable Earth Science (CATEES), to develop a Python software package in order to transition Earth science data processing to the cloud. This project, in particular, supports CATEES and has two primary goals. One, to transition data recipes created by the Goddard Earth Science Data and Information Service Center (GES DISC) into an interactive and educational environment using JupyterNotebooks. Two, to acclimate Earth scientists to cloud computing. To accomplish these goals, we create JupyterNotebooks to compartmentalize the different steps of data analysis and help users obtain and parse data from the command line. We also develop a Docker container, comprised of Jupyter Notebooks, Python dependencies, and command line tools, and configure it into an easy-to-deploy package. The end result is an end-to-end product that simulates the use case of end users working in the cloud computing environment.

discoverability↗

Jupyter-notebook-for-antisymmetrization-circuits

Validation through explicit state-vector validation of the swap operations generated using Dicke-state construction to produce the antisymmetrized states of targets and projectiles for nuclear reaction simulations using quantum computing techniques.

Stetcu, Ionel [Los Alamos National Laboratory]↗

Data From: "Warming and snow loss increase reliance on old groundwater in a Colorado River headwater"

This repository contains the data and code associated with the paper titled "Warming and snow loss increase reliance on old groundwater in a Colorado River headwater," published in Nature Geoscience, 2026. This study seeks to answer how various ages of groundwater interact with mountainous streamflow in mountainous headwaters such as the East River. It includes various model-data processing scripts, primarily for ParFlow-CLM analysis of simulated water years 2015-2021, and two numerical warming experiments (+2.5 and +4.0 degrees C), including run scripts, forcing scripts, and post-processing, as well as comparison to observation datasets, detailed below. This data requires the use of R (.r, .rmd), Python (.py), Jupyter Notebook or Jupyter Lab (.ipynb), ParFLOW-CLM, EcoSLIM. Further information on the use of all file formats mentioned below (e.g. .tff. .nc) are provided within the associated scripts and directory where the files are located. Contents & Usage ASO/: ​​Contains the bash and python scripts used to convert airborne snow observatory (ASO) data (ASO, 2023) in various data formats (georeferenced tiff file, NetCDF, UTM, and to latitude/longitude) then regrided to the ParFlow equivalent grid. Output data are in regrid_regll_data.zip and subsequently visualized and analyzed in plot_and_compare.py for Supplementary Figures A14 and A15. The wksht_ASO_comparison.xlsx spreadsheet is used to calculate the data for Supplementary Figure A16. EcoSLIM/: Contains the scripts and input files to run the EcoSLIM particle tracking simulations (/run_scripts) and the post-processing python script (/plot_scripts/eco_agedist_plots.ipynb). Jasechko et al./: Contains the jupyter notebook (Extract_Elevation.ipynb) to determine the outlet elevations of the 260 watersheds used in Jasechko et al. (2016), and the corresponding table, Table_S1_Watersheds_alt.csv. Used to create Supplementary Information Figure A2. PLM_Wells/: Contains the QA/QC-ed groundwater level time series of the PLM-1 and PLM-6 Monitoring Wells from Faybishenko et al. (2023), reformatted to water years used for Supplementary Figures A19 and and A20. ParFlow/: Contains the input files and run scripts to run ParFlow-CLM (/run_scripts), the python and tool command language (Tcl) scripts to create and distribute the ParFlow forcing simulation files (/forcing), and various scripts and intermediary files to analyze the model outputs (/post_process). SQUIRE/: Contains the processing scripts and intermediary files for the Surface QUantitatIve pRecipitation Estimation (SQUIRE) data (Grover, 2023) used to generate Supplementary Figure A18. USGS_Streamflow/: Contains the raw and gap-filled United States Geological Survey streamflow data (U.S. Geological Survey, 2026) used at the Almont station (site number 09112500). Gap-filling is performed in the R script with data from the Taylor station (site number 09110000). (/USGS_09112500_EAST_RIVER_AT_ALMONT_GAP_FILLED/code_almont_streamflow_gap_fill.Rmd). discharge/: Contains the gap-filled discharge data at the Watershed Function SFA East River pumphouse site (Newcomer et al., 2022) used to generate Supplementary Figure A13 and to compute hourly Nash-Sutcliffe model efficiency coefficients (NSE) in Table A4. snotel_and_flux_tower/: Contains the snow telemetry data (U.S. Department of Agriculture, 2024) from the Butte (site ID 380) and Schofield (site ID 737) stations, reformatted by water year, accessed with the snotelr R package. Used to create Supplementary Figure A17. Also contains the flux tower observational data (FluxTower_Pumphouse_ESS-DIVE.ET_only.h.txt) from Ryken et al. (2022) and sap flux transpiration data (MaxB_Transpiration_5Sites.daily_sums.h.txt) from Ryken (2021), used to create Supplementary Figures A22 and A23, respectively. Raw EcoSLIM model outputs are in excess of 24TB, and are stored on National Energy Research Scientific Computing Center (NERSC) and publicly available via the external link provided in the paper.

atmospheric warming↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 2. Evaluating Controls on Flow Persistence in an Urbanized Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in an urbanized catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, distributed temperature sensing (DTS), continuous self-potential (SP) monitoring, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Field_Application subfolder contains the ATS XML input scripts, data files, output data for the SP site. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. The flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.m can only be used with COMSOL with MATLAB) is executed using the ATS output data to simulate the potential field. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) DTS Contains collated DTS data including raw Stokes and anti-Stokes measurement (provided as .h5 file). It also includes DTS processing.ipynb, a Jupyter notebook for calibrating the DTS data using dts_calibration Python package. cooler_calibration.csv is the DTS calibration CSV used in the calibration sequence. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion. 6) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 7) SP Contains the SP data collected in field at the SP sites (provided as CSV files). 8) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). Note: Code files (.ipynb, .py, .xml) can be opened in any standard code editor, .exo file can be viewed using Paraview, .h5 files can be opened using HDFView software and h5py Python package, and .resipy file can be opened with the open-source ResIPy software.

ATS↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS↗

Data and scripts associated with the manuscript evaluating the hydrologic responses of the Pacific Northwest watersheds to wildfires (v2)

This data package is associated with the publication “Evaluating Post-fire Watershed Response to Varying Burn Severity and Precipitation Regimes Using Fully-distributed and Integrated Hydrologic Models” submitted to Journal of Hydrology (Li et al. 2025). In this study, we employed the Advanced Terrestrial Simulator (ATS), an integrated watershed model that couples surface flow, subsurface flow, and canopy biophysical processes, to investigate post-fire hydrologic responses in a few selected watersheds with varying burn severity.The data package contains the required input data (meteorological forcing, Leaf Area Index, wildfire burn severities, etc.) to run the model, configuration files, the Jupyter notebooks in Python to pre-process and post-process data, the figures in the manuscript, and the modeling output files. The variables include watershed-averaged evapotranspiration, watershed-averaged surface/subsurface/canopy water content, and river discharge at watershed outlet.The data package contains a file-level metadata that lists and describes all the files contained in the data package (ATS_flmd.csv), a data dictionary file that defines columns headers across all csv files contained in the data package (ATS_dd.csv), a data package level readme file (the current file), and four zipped folders.The ‘data’ folder provides data needed to run the model in .h5, .i2s, .xyz, .shp, and .exo formats. The sub-folders are for each data types. The ‘model’ folder provides input files (.xml format) and essential model outputs. Each sub-folder provides the files from each simulated watershed. The ‘notebooks’ folder provides the Jupyter notebooks (.ipynb format) for pre- and post- processing model files, and for producing the figures in the manuscript. The ‘figures’ folder provides the figures associated with manuscript in .pdf and .png formats.The ‘model’ folder and the ‘data’ folder have been split into 5GB-large pieces using the Linux command ‘split -b 5120m model.zip model.zip.’ and ‘split -b 5120m data.zip data.zip.’, respectively. They can be merged back using the Linux command ‘cat model.zip.* > model.zip’ and ‘cat data.zip.* > data.zip’, respectively.

54 ENVIRONMENTAL SCIENCES↗

Connecting Users and Applications with Po.daac Hosted GHRSST Data

The 80+ GHRSST public datasets represent a rich resource for sea surface temperature research and applications given their time series length, resolution, spatial coverage, varying measurement types and processing levels, and availability in the full spectrum of PO.DAAC tools and services ecosystem. The PO.DAAC has created a publicly accessible recipe suite for the user community to perform straightforward yet powerful computations on GHRSST data using python recipes, Jupyter notebooks, R, Matlab, and the NCO programming language. These recipes include numerical computations for regional and global SST trends, anomaly derivations, EOF analysis, climate signal reproduction, and ocean phenology. For example, one recipe reproduces a famous SST based warming figure from the Fourth National Climate Assessment (USA) while another focuses on quantifying the regional changes in ocean SST phenology. Most are python-based while some contain hybrid calls and leverage the NCO programming interface too. All are available on the PO.DAAC user forum (https://podaac.jpl.nasa.gov/forum/) and/or via the open source NASA GitHub repository (https://github.com/nasa/podaac_tools_and_services). Several are available in the Jupyter notebook framework including podaacypy (https://github.com/nasa/podaacpy), a recipe for GHRSST granule metadata discovery and application, and more recently a Jupyter notebook developed to support data analysis and visualization of a cloud-based Zarr formatted Level 4 MUR dataset in the AWS Open Data Registry. Throughout the summer of 2020, the PO.DAAC intends to add and migrate more of its numerical recipes to the Jupyter notebook framework and publish them on its open source GitHub repository.

Gentemann, Chelle↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

A Novel Architecture of JupyterHub on Amazon Elastic Kubernetes Service for Open Data Cube Sandbox

The Open Data Cube (ODC) initiative, with support from the Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has developed a state-of-the-art suite of software tools and products to facilitate the analysis of Earth Observation data. This paper presents a short summary of our novel architecture approach in a project related to the Open Data Cube (ODC) community that provides users with their own ODC sandbox environment. Users can have a sandbox environment all to themselves for the purpose of running Jupyter notebooks that leverage the ODC. This novel architecture layout will remove the necessity of hosting multiple users on a single Jupyter notebook server and provides better management tooling for handling resource usage. In this new layout each user will have their own credentials which will give them access to a personal Jupyter notebook server with access to a fully deployed ODC environment enabling exploration of solutions to problems that can be supported by Earth observation data.

Open Data Cube↗

Data and Scripts associated with “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling.”

This data package is associated with the publication “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling” submitted to Geoscientific Model Development (Muller et al., 2024). In this manuscript, organic matter chemistry and thermodynamics are directly connected to reactive transport simulators through the newly developed Lambda-PFLOTRAN (Parallel Reactive Flow and Transport model) workflow tool that succinctly incorporates organic matter chemistry data generated from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) into reaction networks to simulate aerobic respiration of the organic matter and the resulting biogeochemistry. Lambda-PFLOTRAN is a python-based workflow, executed through a Jupyter Notebook interface, that digests raw FTICR-MS data, develops a representative reaction network based on substrate-explicit thermodynamic modeling (also termed lambda modeling due to its key thermodynamic parameter λ used therein), and completes a biogeochemical simulation with the open source, reactive flow, and transport code PFLOTRAN. This data package contains Jupyter Notebook based workflows for two test cases for running biogeochemical simulations of organic matter oxidation identified by FTICR-MS. It contains four primary folders (workflow, data, src, and analysis), a file-level metadata file (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_flmd.csv) that lists all the files contained in this data package with a short description of each, and a data dictionary (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_dd.csv) file that describes the tabular column headers. The ‘workflow’ folder contains the Jupyter Notebook based workflows for running the lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘data’ folder contains the FTICR-MS data, initial conditions, and incubation data for test cases 1 and 2 in folders titled ‘WHONDRS’ and ‘Colloids’, respectively. The data folder also has a ‘Database’ folder containing a reaction network for bulk organic matter (assumed to be CH2O) and a general database for PFLOTRAN (hanford_rxn_network). The CH2O reaction network defines bulk organic matter oxidation. Biogeochemical simulations are completed for both the lambda binned organic matter and bulk organic matter reaction networks. The ‘hanford_rxn_network’ database includes information required for PFLTORAN simulations including ion size, molar mass, and charge of the aqueous species, gases, and minerals phases. The ‘src’ folder contains python source codes for performing lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘analysis’ folder contains outputs from the test cases 1 and 2 including lambda analysis, PFLOTRAN runs and the calibration results.

54 ENVIRONMENTAL SCIENCES↗