Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Schema Elements for Granta Annual Report: FY2024

Granta: Materials Intelligence (Granta: MI) is a commercial database software distributed by Ansys, Inc. that is utilized by the Nuclear Security Enterprise (NSE) to organize and store relevant materials data. Lack of standard and well-documented database schema is the primary obstacle to an NSE materials data management solution, so the objective of this project is to create and document such a schema. In FY21, an approach for designing, documenting, and managing a standard database schema was described based on the creation of schema elements (collections of attributes used to describe particular aspects of the data) to be used as building blocks for creating various database tables without duplication. In FY22, these methods were applied through a multi-site collaboration to create and document the schema elements necessary to build a thermogravimetric analysis (TGA) testing table. In FY23 the schema was expanded to include elements for a differential scanning calorimetry (DSC) table, along with schema for supporting metadata tables including Instruments, Projects, Documents, and Testing Series. In FY24 the following progress was made, again through multi-site collaboration: • The existing schema elements were modified to accommodate thermomechanical analysis (TMA) data, and a table, Test Data: TMA, was created for managing TMA data. • The elements necessary for the following additive manufacturing (AM) data tables (directed at data specific to selective laser sintering AM technology) were created: • AM Builds • AM Processes • AM Part Designs • Built AM Parts • AM Feedstock Materials • AM Feedstock Material Batches • The elements necessary for creating a Calibrated Material Models table were created, and the Calibrated Material Models table was created. In FY25 the existing schema will be deployed on the production enterprise Granta instance on the enterprise secure network. Schema elements will be appended, and new elements created as necessary, to allow the creation of tables specifically to support materials testing, AM process development, and design and analysis for modernization programs.

36 MATERIALS SCIENCE↗

Vegetation Warming Experiment: Landscape-scale digital camera imagery for vegetation phenology, Utqiagvik (Barrow), Alaska, 2018

Images captured using a StarDot NetCam SC phenocamera looking east from the top of the Barrow Environmental Observatory (BEO) Sled Shed, Utqiagvik, Alaska. The camera was installed to remotely monitor plant phenology and operation of the BNL TEST group's ZPW (Zero Power Warming) chambers during the growing season of 2018. Images were captured from early spring (13 April) through to mid fall (18 October). Snowmelt, vegetation growth and senescence, and snow accumulation were captured. Images were uploaded to the BNL FTP server every hour until 16 June, and then every 10 minutes. Files were renamed with the date and time of the image. jpg images have been compressed in *.tar.gz format (7.9 GB). Different fields of view (northeasterly) were also captured using 3 Wingscapes TimelapseCam cameras mounted on a mast on the sled shed. Images were recorded from 21 June to 24 September, 2018, at 30 minute intervals, from 11:00 - 14:30, Alaska daylight time (AKDT, UTC-8). Movie files (.mp4) of each Wingscapes camera dataset. Additional metadata included in *.pdf and *.csv files. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Vegetation Warming Experiment: Chamber and ambient plot digital camera imagery for vegetation phenology, Utqiagvik (Barrow), Alaska, 2017

Time lapse photography (*.jpg) of experimental plots within five warming chambers (ZPWs) and paired control plots located on the Barrow Environmental Observatory (BEO), Barrow (now Utqiagvik), Alaska. Images were recorded from 21 June to 18 September, 2017 to capture vegetation dynamics during the growing season. Vegetation phenology, including green up and senescence were captured. Images were recorded daily at 30 minute intervals, from 11:00-14:30 Alaska daylight time (AKDT, UTC-8), using Wingscapes TimelapseCam cameras. The target species was Petasites frigidus. Also included multiple metadata files including reporting formats as *.pdf and *.csv files. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

The Zooplankton International Geospatial dataset: A global repository of spatiotemporal freshwater zooplankton community composition data from lakes and reservoirs to support ecological research

Zooplankton transfer substantial energy in aquatic food webs and are used as indicators of environmental change. Syntheses of zooplankton community dynamics globally require datasets that span a wide range of environmental gradients; however, these datasets are limited due to methodological differences across programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we created the Zooplankton International Geospatial (ZIG) dataset, which includes original zooplankton, water physical and chemical variables, and lake morphometric data from 311 inland lakes and reservoirs. ZIG includes waterbodies ranging in size from 0.005 to 82,100 km2 and spanning broad latitudinal (−47.26 to 64.90) and longitudinal ranges (−165.04 to 176.53). Temporal coverage for individual waterbodies ranges between 1 and 60 yr with sampling frequency ranging from annually to weekly. With its extensive coverage and content, we consider ZIG to be a cornerstone for future investigations of global scale lake biodiversity change.

Figary, Stephanie [Cornell University, Ithaca, NY]↗

Proximal remote sensing: an essential tool for bridging the gap between high‐resolution ecosystem monitoring and global ecology

Summary A new proliferation of optical instruments that can be attached to towers over or within ecosystems, or ‘proximal’ remote sensing, enables a comprehensive characterization of terrestrial ecosystem structure, function, and fluxes of energy, water, and carbon. Proximal remote sensing can bridge the gap between individual plants, site‐level eddy‐covariance fluxes, and airborne and spaceborne remote sensing by providing continuous data at a high‐spatiotemporal resolution. Here, we review recent advances in proximal remote sensing for improving our mechanistic understanding of plant and ecosystem processes, model development, and validation of current and upcoming satellite missions. We provide current best practices for data availability and metadata for proximal remote sensing: spectral reflectance, solar‐induced fluorescence, thermal infrared radiation, microwave backscatter, and LiDAR. Our paper outlines the steps necessary for making these data streams more widespread, accessible, interoperable, and information‐rich, enabling us to address key ecological questions unanswerable from space‐based observations alone and, ultimately, to demonstrate the feasibility of these technologies to address critical questions in local and global ecology.

Plant Sciences↗

Data and scripts associated with “Point-scale organic-matter decomposition in streambeds is weakly associated with reach-scale respiration”

This data package is associated with “Point-scale organic-matter decomposition in streambeds is weakly associated with reach-scale respiration” published in EGU Biogeosciences (Stegen et al., 2026; https://doi.org/10.5194/bg-23-3981-2026). It contains cotton strip decomposition rates (Kcd and Kdd) collected across the Yakima River Basin (YRB), Washington, USA. These data were collected to support a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin. Associated data used in analysis, metadata, and field protocols can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1969566, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1987520. This data package is associated with the repository found at https://github.com/river-corridors-sfa/rcsfa-ST-2B-SSS-cotton-strip. A preliminary version of this data package was published in December 2025 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This data package consists of (1) readme; (2) data dictionary (dd); (3) file level metadata (flmd); and (4) four folders: (1) R-scripts; (2) figures; (3) outputs from the scripts; and (4) published data. The published data folder contains a readme directing the user to download data in order to run the R-scripts. All files are .csv, .pdf, .R, .Rmd, and .txt. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY↗

Data from a four-day long microcosm experiment addressing the destabilization of artificial mineral-associated organic matter by model root exudates embedded in a soil matrix from the Rocky Mountain Biological Laboratory (Gothic, CO, USA), 2019

This dataset provides data collected during a four-day long laboratory soil microcosm experiment testing the efficacy of root exudate-driven mineral-associated organic matter destabilization. This dataset contains four data files in comma-separate values (*.csv). The files provide the metadata and the experimental results on microbial respiration, MAOM-derived respiration, and sequential mineral-extractions. This data was used to produce the figures in Bölscher et al., 2026. The results of the experiment can be found in the open access article Bölscher et al., 2026 (https://doi.org/10.1016/j.soilbio.2026.110276). Abstract: Mineral-associated organic matter (MAOM) is often considered stable, but root exudates can destabilize MAOM via various pathways. Theory and model system studies suggest that direct MAOM destabilization by strong ligands, like oxalic acid, or reducing agents, like catechol, is more effective than indirect, microbial-mediated MAOM destabilization, stimulated by less reactive compounds like glucose. Here, we demonstrate that the presence of a soil matrix alters the efficacy of exudate-driven MAOM destabilization pathways. Glucose and catechol destabilized significantly greater amounts of MAOM from ferrihydrite and aluminum hydroxide (Al (OH)3) embedded in a soil matrix than oxalic acid. Our findings indicate that indirect, microbial-mediated MAOM destabilization may play a larger role than direct MAOM destabilization in soil environments.

Destabilization↗

CT Scans of Cores Metadata, Utqiagvik (Barrow), Alaska, 2015

Individual ice cores were collected from Barrow Environmental Observatory in Barrow, Alaska, throughout 2013 and 2014. Cores were drilled along different transects to sample polygonal features (i.e. the trough, center and rim of high, transitional and low center polygons). Most cores were drilled around 1 meter in depth and a few deep cores were drilled around 3 meters in depth. Three-dimensional images of the frozen cores were constructed using a medical X-ray computed tomography (CT) scanner. TIFF files can be uploaded to ImageJ (an open-source imaging software) to examine soil structure and soil densities within each core.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Closure Report for Corrective Action Unit 116: Area 25 Test Cell C Facility, Nevada National Security Site, Nevada with ROTC-1

CR loaded to this OSTI record. Just adding new file which includes CR plus new ROTC 1 and update the metadata to the following: This Closure Report (CR) presents information supporting closure of Corrective Action Unit (CAU) 116, Area 25 Test Cell C Facility. This CR complies with the requirements of the Federal Facility Agreement and Consent Order (FFACO) that was agreed to by the State of Nevada; the U.S. Department of Energy (DOE), Environmental Management; the U.S. Department of Defense; and DOE, Legacy Management (FFACO, 1996 [as amended March 2010]). CAU 116 consists of the following two Corrective Action Sites (CASs), located in Area 25 of the Nevada National Security Site: (1) CAS 25-23-20, Nuclear Furnace Piping and (2) CAS 25-41-05, Test Cell C Facility. CAS 25-41-05 consisted of Building 3210 and the attached concrete shield wall. CAS 25-23-20 consisted of the nuclear furnace piping and tanks. Closure activities began in January 2007 and were completed in August 2011. Activities were conducted according to Revision 1 of the Streamlined Approach for Environmental Restoration Plan for CAU 116 (U.S. Department of Energy, National Nuclear Security Administration Nevada Site Office [NNSA/NSO], 2008). This CR provides documentation supporting the completed corrective actions and provides data confirming that closure objectives for CAU 116 were met. Site characterization data and process knowledge indicated that surface areas were radiologically contaminated above release limits and that regulated and/or hazardous wastes were present in the facility. The Record of Technical Change 1 updated the use restriction information.

54 ENVIRONMENTAL SCIENCES↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

Reproductive and leaf litterfall fluxes in forest ecosystem sites globally (1950-2022)

Forest allocation of net primary productivity (NPP) to reproduction is poorly quantified globally, despite its critical role in forest regeneration and a well-supported trade-off with allocation to growth. Although field measurements of total NPP are rare, our work finds that a proxy for reproductive carbon allocation constructed from leaf (L) and reproductive (R) litterfall fluxes, R/(R+L), is strongly correlated with R/NPP, facilitating analysis across a wide range of sites where biometric estimates of NPP are not available (R² = 0.85; Hanbury-Brown et al., 2022, Ward et al., in prep). To investigate relationships between ecosystem-scale reproductive allocation (RA) and climate, soil fertility, and stand age gradients, we conducted a literature search and synthesized 824 observations of annual average leaf and reproductive litterfall fluxes across forest sites globally. The zip file includes 1) a folder Data/ containing the litterfall data ("GlobalForestRA_data.csv") and metadata ("GlobalForestRA_metadata.doc") files. The data file includes geographic coordinates, long-term mean annual temperature and precipitation (1970-2000, extracted from WorldClim2.1), leaf and reproductive litterfall fluxes, sampling interval and protocols, forest characteristics (dominant leaf morphology, information pertaining to forest age and successional stage, and disturbance history) and soil properties (% sand, %silt, %clay, total phosphorus (P), nitrogen (N), cation exchange capacity (CEC) and pH) extracted from SoilGrids250 and from on-site measurements, where available. The metadata file contains information about each variable reported in the data file, including data sources, processing methods, and all references. The Data folder contains two additional files used to create Figure 1; these are described in greater detail in the README.2) R scripts GloalForestRA_analysis.r and GlobalForestRA_SI.r and a folder /Functions used to produce results, figures, and tables in the manuscript Ward et al. (in press)3) a README file describing how the data and R scripts can be used to reproduce statistical results, figures, and tables found in the manuscript. Ward et al. (in press)This repository can also be found at: https://github.com/r-ward/Global_Analysis_ForestRA.Ward, R.E., Zhang-Zheng, H. Aernethy, K., Adu-Bredu, S., Arroyo, L., Bailey, A. et al. (in press). Forest age rivals climate to explain reproductive allocation patterns in forest ecosystems globally. Ecology Letters. Hanbury-Brown, A.R., Ward, R.E. & Kueppers, L.M. (2022). Forest regeneration within Earth system models: current process representations and ways forward. New Phytol., 235, 20–40.Ward et al. (2025), Forest age rivals climate to explain reproductive allocation patterns in forest ecosystems globally, in prep.

54 ENVIRONMENTAL SCIENCES↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Rhodotorula toruloides Nitrogen Limitation PTM Profiling Multi-Omics (TZ-DP1)

The purpose of this experiment was to evaluate the regulatory stress response of Oleaginous yeast species Rhodotorula toruloides NBRC 0880 (JGI strain IFO0880 v4.0) under nitrogen-rich and nitrogen-limited conditions over time. Time course experimental samples (0, 24, 48, and 72 hours after inoculation) were prepared using a semi-automated multi-PTM proteomic approach, using tandem mass tag 18-plex (TMT18), and lipidome remodeling for downstream multi-omics analysis. Processed datasets are openly accessible from PNNL DataHub and contain secondary processed proteomic (redox, phospho, and global TMT) and lipidomic (positive and negative ion mode) results files and experimental design metadata.

59 BASIC BIOLOGICAL SCIENCES↗