Search NASA⌕ Search

SEARCH · Search NASA

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Physics-informed machine learning for building performance simulation-A review of a nascent field

Building performance simulation (BPS) is critical for understanding building dynamics and behavior, analyzing the performance of the built environment, optimizing energy efficiency, improving demand flexibility, and enhancing building resilience. However, conducting BPS is not trivial. Traditional BPS relies on accurate building energy models, which are primarily physics-based and heavily dependent on detailed building information, expert knowledge, and case-by-case model calibrations, significantly limiting their scalability. With the development of sensing technology and the increased availability of data, there is growing attention and interest in data-driven BPS. However, purely data-driven models often suffer from limited generalization ability and a lack of physical consistency, resulting in poor performance in real-world applications. To address these limitations, recent studies have begun integrating physics priors into data-driven models, a methodology known as physics-informed machine learning (PIML). PIML is an emerging field where its definitions, methodologies, evaluation criteria, application scenarios, and future directions remain open. To bridge those gaps, this study systematically reviews the state-of-the-art PIML for BPS, offering a comprehensive definition of PIML and comparing it to traditional BPS approaches regarding data requirements, modeling effort, performance, and computational cost. We also summarize the commonly used methodologies, validation approaches, application domains, available data sources, open-source packages, and testbeds. In addition, this study provides a general guideline for selecting appropriate PIML models based on BPS applications. Finally, this study identifies key challenges and outlines future research directions, providing a solid foundation and valuable insights to advance R&D of PIML in BPS.

Jiang, Zixin↗

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

1000 Soils Pilot Dataset, version 8, May 2025

This record hosts data generated by the 1000 Soils Pilot. Data will be updated as more become available. Please see the most recent data upload for current data. A beta visualization tool is available for some data types at https://shinyproxy.emsl.pnnl.gov/app/1000soils. Please submit any suggestions or comments through the 'contact' tab. We are actively working to improve visualizations and value all feedback. Data completed include: Geochemistry, texture, respiration, and enzyme activities FTICR-MS organic matter chemistry Microbial biomass C and N TOC/TDN of water-extractable OM X-ray computed tomography (derived metrics available here, raw data available upon request) Metagenomes; a variety of data formats are available upon request Soil hydraulic properties Data in progress: LC-MS/MS in development, timeline TBD, inquire for status 1000S_processed_BGC_summary.csv contains all available biogeochemical data; microbial biomass C and N; and TOC/TDN of water-extractable OM; and 1000S_Tomography.xslx contains a summary of data generated via X-ray computed tomography. icr_v2_corems2.csv contains FTICR-MS data processed by CoreMS version 2. These data are merged by formula across instrument runs to enable cross-sample comparisons. Technical replicates are merged by retaining peaks present in 2 out of 3 replicates. 1000Soils_Metadata_Site_Mastersheet_v1.csv contains site information. Soil Hydraulics_corrected_02042025.xlsx contains soil hydraulics information. Readme File_v4.xlsx is the readme file. Please contact the MONet project (monet.emsl@pnnl.gov) or Emily Graham (emily.graham@pnnl.gov) with questions. The following file and all raw data are available upon request: icr_by_mass_for_single_sample_analysis_only.csv contains FTICR-MS data processed by CoreMS and is intended for usage in the calculation of biochemical transformations within samples only. These data are not acceptable for cross-sample comparison of masses because they are from multiple instrument runs. For more information, please see: https://www.emsl.pnnl.gov/monet and https://sc-data.emsl.pnnl.gov/monet Acknowledgment: Soil data were provided by the Molecular Observation Network (MONet) at the Environmental Molecular Sciences Laboratory (https://ror.org/04rc0xn13), a DOE Office of Science user facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. The work (proposal: 10.46936/10.25585/60008970) conducted by the U.S. Department of Energy, Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science user facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231. The Molecular Observation Network (MONet) database is an open, FAIR, and publicly available compilation of the molecular and microstructural properties of soil. Data in the MONet open science database can be found at https://sc-data.emsl.pnnl.gov/.

biogeochemistry↗

Evaluating grid stress and reliability in future electricity grids across a range of demand, generation mix, and weather trends

The reliability of power grids in the future will depend on how system planners account for the integration of new technologies, extreme weather events, and uncertainties in demand growth from increased electrification and data centers. This study introduces an open-source, multisectoral, multiscale modeling framework that projects grid stress and reliability trends between 2020 and 2055 in the Western Interconnection of the United States. The framework integrates global to national energy-water-land dynamics with power plant siting and hourly grid operations modeling. We analyze future wholesale electricity price shocks and unserved energy events across eight scenarios spanning a range of population growth and economic change, generation mixes, and weather conditions. Our results show future grids with high percentage of non-renewable generation and strong economic growth are characterized by higher reliability and lower wholesale electricity prices than lower growth scenarios because of larger reliance on dispatchable generators and lower fossil fuel extraction costs. Scenarios with high percentage of renewable resources have lower median but more volatile wholesale electricity prices as well as more frequent and severe unserved energy events compared to scenarios relying more on dispatchable generators. These events occur because higher proportion of solar and wind energy causes net demand curves to deepen during midday (duck curves get progressively severe), exacerbating the challenge of meeting demand during summer evening peaks. This study suggests that robust and co-optimized transmission and energy storage planning could help maintain low wholesale electricity prices and high reliability levels in future electricity grids across uncertainties in generation mixes.

Electric grid reliability↗

Thermodynamics and its prediction and CALPHAD modeling: Review, state of the art, and perspectives

Thermodynamics is a science concerning the state of a system, whether it is stable, metastable, or unstable, when interacting with its surroundings. The combined law of thermodynamics derived by Gibbs about 150 years ago laid the foundation of thermodynamics. In Gibbs combined law, the entropy production due to internal processes was not included, and the 2nd law was thus practically removed from the Gibbs combined law, so it is only applicable to systems under equilibrium, thus commonly termed as equilibrium or Gibbs thermodynamics. Gibbs further derived the classical statistical thermodynamics in terms of the probability of configurations in a system in the later 1800's and early 1900's. With the quantum mechanics (QM) developed in 1920's, the QM-based statistical thermodynamics was established and connected to classical statistical thermodynamics at the classical limit as shown by Landau in the 1940's. In 1960's the development of density functional theory (DFT) by Kohn and co-workers enabled the QM prediction of properties of the ground state of a system. On the other hand, the entropy production due to internal processes in non-equilibrium systems was studied separately by Onsager in 1930's and Prigogine and co-workers in the 1950's. In 1960's to 1970's the digitization of thermodynamics was developed by Kaufman in the framework of the CALculation of PHAse Diagrams (CALPHAD) modeling of individual phases with internal degrees of freedom. CALPHAD modeling of thermodynamics and atomic transport properties has enabled computational design of complex materials in the last 50 years. Our recently termed zentropy theory integrates DFT and statistical mechanics through the replacement of the internal energy of each individual configuration by its DFT-predicted free energy. The zentropy theory is capable of accurately predicting the free energy of individual phases, transition temperatures and properties of magnetic and ferroelectric materials with free energies of individual configurations solely from DFT-based calculations and without fitting parameters, and is being tested for other phenomena including superconductivity, quantum criticality, and black holes. Those predictions include the singularity at critical points with divergence of physical properties, negative thermal expansion, and the strongly correlated physics. Furthermore, those individual configurations may thus be considered as the genomic building blocks of individual phases in the spirit of the materials genome®. This has the potential to shift the paradigm of CALPHAD modeling from being heavily dependent on experimental inputs to becoming fully predictive with inputs solely from DFT-based calculations and machine learning models built on those calculations and existing experimental data through newly developed and future open-source tools. Furthermore, through the combined law of thermodynamics including the internal entropy production, it is shown that the kinetic coefficient matrix of independent internal processes is diagonal with respect to the conjugate potentials in the combined law, and the cross phenomena that the phenomenological Onsager flux and reciprocal relationships are due to the dependence of the conjugate potential of a molar quantity on nonconjugate molar quantities and other potentials, which can be predicted by the zentropy theory and CALPHAD modeling.

42 ENGINEERING↗

WellPINN: Accurate Well Representation for Transient Fluid Pressure Diffusion in Subsurface Reservoirs With Physics‐Informed Neural Networks

Accurate representation of pumping wells is essential for reliable reservoir characterization and simulation of operational scenarios in subsurface flow models. Physics-informed neural networks (PINNs) are emerging as a promising alternative to numerical models for reservoir modeling, offering seamless integration of monitoring data and governing physical equations. However, existing PINN-based studies face major challenges in capturing fluid pressure near wells when using a source/sink term, particularly during the early stages after pumping begins. We address this problem by introducing WellPINN, a workflow in which an initially trained PINN infers fluid pressure across the entire reservoir domain using a large equivalent well radius. This initial PINN solution is then locally refined around the well by a set of subdomain PINNs that are trained for smaller equivalent well radii. Continuity across these subdomain interfaces as well as at the initial condition is ensured by hard-constraining each PINN on its subdomain boundary. Our results demonstrate WellPINN as the first workflow of its kind to focus on accurate inference of fluid pressure from pumping rates throughout the entire injection period, significantly advancing the potential of PINNs for inverse modeling and operational scenario simulations. All data and code for this paper are openly available at https://doi.org/10.20350/DIGITALCSIC/17260.

58 GEOSCIENCES↗

Vibrational and electronic properties of Np 2 O 5 from experimental spectroscopy and first principles calculations

High-valence actinide oxides are critical to understanding the behavior of 5f-electrons, yet their structural and electronic properties remain poorly understood due to challenges in synthesis and handling. We report the first Raman spectroscopic study of single-crystalline Np 2 O 5 and the first scanning tunneling spectroscopy (STS) measurement on any neptunium-containing material. Hydrothermally synthesized crystals were structurally verified by X-ray diffraction. Raman spectra revealed sharply resolved vibrational features, including previously unreported low-frequency modes. STS measurements revealed a band gap of 1.5 eV. Density functional theory (DFT) enables vibrational mode assignments, reveals neptunium-dominated low-frequency phonons, oxygen-dominated high-frequency modes, and predicts an indirect band gap of 1.68 eV. This predicted value is in excellent agreement with the experimentally measured STS gap. This combined Raman, DFT, and STS approach provides a robust framework for correlating lattice dynamics and electronic structure in actinide materials, providing benchmark data for Np 2 O 5 , and opening new avenues for probing structure–property relationships in complex f-electron materials.

36 - MATERIALS SCIENCE↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

QEpsilon v0.1.0

QEpsilon is a Python package designed to minimize the effort required to build a data-driven quantum master equation of an open quantum system and to perform time evolution of the master equation. Applications of QEpsilon span from quantum computing to condensed matter systems.

Xie, Pinchen [Lawrence Berkeley National Laborator↗

Coupling of high-resolution mass spectrometer and photosynthesis system for comprehensive leaf volatile metabolite profiling

Background Leaf-level biogenic volatile organic compounds (BVOCs) emissions represent a major source of organic gases in the atmosphere, influencing both climate and air quality. These emissions are strongly driven by environmental perturbations, which affect individual plant- to ecosystem-level processes. Uncovering all the BVOCs and understanding how their emissions respond to altered environmental conditions provide critical insights into vegetation-driven changes in atmospheric chemistry. We developed a tandem instrumentation setup that integrates a proton transfer reaction time-of-flight mass spectrometer (PTR-ToF-MS) with parts-per-trillion detection limits and a photosynthetic infrared gas exchange system for the untargeted survey of all the BVOCs. This novel system enables simultaneous, real-time monitoring of BVOC emissions and photosynthetic parameters at the leaf level, offering new opportunities to disentangle the physiological and environmental drivers of VOC release. Furthermore, we established the VOC Analysis and Processing Optimization Resource (VAPOR), an open-access software tool designed for rapid data post-processing and the analysis of the variability of hundreds of BVOCs. We assessed the performance of the tandem system under varying background conditions, using standard gas mixtures and a range of environmental factors. Results Blank emissions were substantially lower for major BVOCs (e.g., isoprene) compared to those observed in plant emissions. Despite this, the observation of background-level VOCs highlights the importance of routinely acquiring and accounting for blank measurements in analyses using the coupled instrumentation. Introduction of known VOC concentrations to the system demonstrated a linear response across different compounds with varying molecular compositions, indicating minimal gas loss regardless of chemical moieties within the coupled instrumentation. We applied the optimized system to investigate the physiological mechanisms driving BVOC emissions across different genotypes of poplar and pennycress. The high mass resolution capabilities of the PTR-ToF-MS, coupled with comprehensive VAPOR-driven data analysis, enabled the identification of several important BVOCs, including methanol and methanethiol; these BVOCs displayed substantial variation across pennycress genotypes and showed concentrations ~ 100–350% higher than the blank. Moreover, isoprene emissions varied significantly among poplar genotypes grown in different potting media. Conclusions Tandem instrumentation offers a powerful tool for profiling volatile molecular markers and elucidating their genetic and environmental underpinnings. This approach enhances our ability to predict BVOC emissions in response to genotype by environmental interactions and contributes to a deeper understanding of vegetation responses to environmental changes.

Biogenic volatile organic compounds↗

High Resolution Siting Suitability of Various Power Plant Technologies

Energy sector planning models determine the aggregate need for new generation, but these models are typically at the state or regional scale and are not equipped to address the wide range of location- and technology-specific issues that are increasingly a factor in power plant siting. These animations demonstrate the aggregate siting suitability of various power plant technology configurations, considering technology-specific factors that can prohibit development. The data presented is from the GRIDCERF (Geospatial Raster Input Data for Capacity Expansion Regional Feasibility) data package. GRIDCERF is a harmonized, open-source geospatial product that can be used to evaluate siting suitability for renewable and non-renewable power plants in the conterminous United States. The animations presented here demonstrate a curated selection of the full suite of technology configurations available. GRIDCERF provides the necessary inputs for models that simulate power plant siting for regional capacity expansion planning such as the Capacity Expansion Regional Feasibility (CERF) model.

Mongird, Kendall [Pacific Northwest National Labor↗

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS↗

A Survey of Open Source Software Repositories in the U.S. Department of Energy’s National Laboratories

There are 17 national laboratory systems in the United States operating under the auspices of the U.S. Department of Energy (DOE). These government labs employ tens of thousands of people engaging in research software engineering activities across a variety of missions. To support this work, many open source projects are maintained. Further, many of these projects have broad utility to the computing community at large and domain scientists in a variety of fields. However, the complexity and decentralized nature of the laboratory system has resulted in a situation where no one entity even knows about all the open source software projects in this ecosystem, let alone crude metrics of their health. In this article, we do the first external inventory of open source software repositories with a nexus to DOE labs. We posit that a project’s need for sustainability support can be determined by comparing measures of active use to measures of active maintenance.

97 MATHEMATICS AND COMPUTING↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗