Search NASA⌕ Search

SEARCH · Search NASA

Results for “molecular data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unsupervised learning of representative local atomic arrangements in molecular dynamics data

Molecular dynamics (MD) simulations present a data-mining challenge, given that they can generate a considerable amount of data but often rely on limited or biased human interpretation to examine their information content. By not asking the right questions of MD data we may miss critical information hidden within it. Here we combine dimensionality reduction (UMAP) and unsupervised hierarchical clustering (HDBSCAN) to quantitatively characterize prevalent coordination environments of chemical species within MD data. By focusing on local coordination, we significantly reduce the amount of data to be analyzed by extracting all distinct molecular formulas within a given coordination sphere. We then efficiently combine UMAP and HDBSCAN with alignment or shape-matching algorithms to partition these formulas into structural isomer families indicating their relative populations. The method was employed to reveal details of cation coordination in electrolytes based on molecular liquids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

Towards a unified nonlocal, peridynamics framework for the coarse-graining of molecular dynamics data with fractures

Molecular dynamics (MD) has served as a powerful tool for designing materials with reduced reliance on laboratory testing. However, the use of MD directly to treat the deformation and failure of materials at the mesoscale is still largely beyond reach. In this work, we propose a learning framework to extract a peridynamics model as a mesoscale continuum surrogate from MD simulated material fracture data sets. Firstly, we develop a novel coarse-graining method, to automatically handle the material fracture and its corresponding discontinuities in the MD displacement data sets. Inspired by the weighted essentially non-oscillatory (WENO) scheme, the key idea lies at an adaptive procedure to automatically choose the locally smoothest stencil, then reconstruct the coarse-grained material displacement field as the piecewise smooth solutions containing discontinuities. Then, based on the coarse-grained MD data, a two-phase optimization-based learning approach is proposed to infer the optimal peridynamics model with damage criterion. In the first phase, we identify the optimal nonlocal kernel function from the data sets without material damage to capture the material stiffness properties. Then, in the second phase, the material damage criterion is learnt as a smoothed step function from the data with fractures. As a result, a peridynamics surrogate is obtained. As a continuum model, our peridynamics surrogate model can be employed in further prediction tasks with different grid resolutions from training, and hence allows for substantial reductions in computational cost compared with MD. We illustrate the efficacy of the proposed approach with several numerical tests for the dynamic crack propagation problem in a single-layer graphene. Our tests show that the proposed data-driven model is robust and generalizable, in the sense that it is capable of modeling the initialization and growth of fractures under discretization and loading settings that are different from the ones used during training.

97 MATHEMATICS AND COMPUTING↗

Molecular simulation data for 'Data-guided Multi-Map variables for ensemble refinement of molecular movies'

These trajectories, scripts, and analysis performed on Summit underly the work published as 'Data-guided Multi-Map variables for ensemble refinement of molecular movies'. The trajectories include equilibrium and non-equilibrium sampling of ADK, CODH, and FLPP3, the scripts used to build the systems, and the scripts used to analyze the output. The directory structure is explained further in an internal README file.

59 BASIC BIOLOGICAL SCIENCES↗

Effects of 9.5 Years of Whole-Soil Warming on the Fatty Acid and n-Alkanes Composition in Bulk Soil and Density Fractions at Blodgett Experimental Forest, California, USA

Original data of molecular data (fatty acids and n-alkanes) including concentrations and calculated molecular proxies in a whole-soil warming experiment at the Blodgett Forest Research Station after 9.5 years of warming. The study site has a Mediterranean climate with annual average temperature of 12.5 ℃ and annual average precipitation of 1774 mm. The study site is characterized by a mesic Ultic Alfisol formed from granitic parent material, corresponding to a Dystric Cambisol under the World Reference Base for Soil Resources (WRB) classification system. Experimental warming is applied throughout the soil profile to a depth of 1 m using vertically embedded heating cables that raise soil temperature by 4 °C relative to ambient conditions. Soil samples were collected on 1 May 2023, after the experiment had been operating continuously for about 9.5 years since its initiation in January 2014.The data has been processed from raw data and cross-validated by other peers. The dataset includes: - Bulk_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including Carbon Preference Index (CPI) and Average Chain Length (ACL) of bulk soil organic carbon; - Fractions_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including CPI and ACL of free particulate organic matter (fPOM) and mineral-associated organic matter (MAOM); - Bulk_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of bulk soil organic carbon; - Fractions_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of fPOM and MAOM; - n-Alkanes_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the n-alkane monomers identified and integrated for bulk soil, fPOM and MAOM; - Fattyacid_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the fatty acid monomers including diacids identified and integrated for bulk soil, fPOM, and MAOM. All data are provided in CSV format and can be viewed using Microsoft Excel. We specifically look at fatty acids (FA) and n-alkanes in bulk soil, fPOM and MAOM and calculated molecular proxies such as CPI and ACL to understand the source of oragnic carbon (with ACL) and degree of decomposition (CPI) of each soil fraction. Due to lack of long-chain fatty acids (carbon number ⩾ 20), microorganism-derived organic carbon is characterized by shorter ACL in comparison to plant-derived organic carbon. Fresh SOC is characterized by even-over-odd dominance for fatty acids and odd-over-even dominance for n-alkanes. Therefore, CPI indicates whether soil organic carbon (SOC) represents fresh input (CPI > 10) or is strongly decomposed (close to 1). The research questions should be then, after 9.5-year warming: 1. whether the relative contribution between microorganism-derived and plant-derived SOC in each soil fraction? 2. whether fPOM became more decomposed whereas MAOM remained relatively persistent in each soil fraction across the soil depth?

Carbon↗

Neutral Loss Mass Spectral Data Enhances Molecular Similarity Analysis in $\mathrm{METLIN}$

We report Neutral loss (NL) spectral data presents a mirror of MS 2 data and is a valuable yet largely untapped resource for molecular discovery and similarity analysis. Tandem mass spectrometry (MS 2 ) data is effective for the identification of known molecules and the putative identification of novel, previously uncharacterized molecules (unknowns). Yet, MS 2 data alone is limited in characterizing structurally related molecules. To facilitate unknown identification and complement the METLIN-MS 2 fragment ion database for characterizing structurally related molecules, we have created a MS 2 to NL converter as a part of the METLIN platform. The converter has been used to transform METLIN’s MS 2 data into a neutral loss database (METLIN-NL) on over 860 000 individual molecular standards. The platform includes both the MS 2 to NL converter and a graphical user interface enabling comparative analyses between MS 2 and NL data. Examples of NL spectral data are shown with oxylipin analogues and two structurally related statin molecules to demonstrate NL spectra and their ability to help characterize structural similarity. Mirroring MS 2 data to generate NL spectral data offers a unique dimension for chemical and metabolite structure characterization.

59 BASIC BIOLOGICAL SCIENCES↗

On the integration of molecular dynamics, data science, and experiments for studying solvent effects on catalysis

Computational workflows that combine molecular dynamics (MD) simulations and emerging data-centric (DC) methods can accelerate the screening and analysis of solvent systems experimentally and computationally. Here, MD simulations provide atomic positions and velocities of reactant, solvent, and catalyst materials that can be manipulated into data representations that in turn can be used by DC techniques to conduct predictive modeling, feature extraction, and experimental design. For liquid-phase catalytic applications, emerging DC techniques such as Convolutional and Graph Neural Networks (CNN/GNN), Topological Data Analysis (TDA), and Active Learning (AL) can leverage MD and experimental data to quickly predict solvent effects on reaction outcomes. For instance, in recent studies, 3D solvent environments obtained with MD have been exploited by CNNs to predict experimental reaction rates for homogeneous acid-catalyzed lignocellulosic processes. In this perspective, we discuss basic principles of DC methods and how these can be combined with MD to enable high-throughput screening of solvent selection for diverse catalysis applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Morphological Characters Can Strongly Influence Early Animal Relationships Inferred from Phylogenomic Data Sets

There are considerable phylogenetic incongruencies between morphological and phylogenomic data for the deep evolution of animals. This has contributed to a heated debate over the earliest-branching lineage of the animal kingdom: the sister to all other Metazoa (SOM). Here, we use published phylogenomic data sets ($\sim $45,000–400,000 characters in size with $\sim $15–100 taxa) that focus on early metazoan phylogeny to evaluate the impact of incorporating morphological data sets ($\sim $15–275 characters). We additionally use small exemplar data sets to quantify how increased taxon sampling can help stabilize phylogenetic inferences. We apply a plethora of common methods, that is, likelihood models and their “equivalent” under parsimony: character weighting schemes. Our results are at odds with the typical view of phylogenomics, that is, that genomic-scale data sets will swamp out inferences from morphological data. Instead, weighting morphological data 2–10$\times $ in both likelihood and parsimony can in some cases “flip” which phylum is inferred to be the SOM. This typically results in the molecular hypothesis of Ctenophora as the SOM flipping to Porifera (or occasionally Placozoa). However, greater taxon sampling improves phylogenetic stability, with some of the larger molecular data sets ($>$200,000 characters and up to $\sim $100 taxa) showing node stability even with $\geqq100\times $ upweighting of morphological data. Accordingly, our analyses have three strong messages. 1) The assumption that genomic data will automatically “swamp out” morphological data is not always true for the SOM question. Morphological data have a strong influence in our analyses of combined data sets, even when outnumbered thousands of times by molecular data. Morphology therefore should not be counted out a priori. 2) We here quantify for the first time how the stability of the SOM node improves for several genomic data sets when the taxon sampling is increased. 3) The patterns of “flipping points” (i.e., the weighting of morphological data it takes to change the inferred SOM) carry information about the phylogenetic stability of matrices. The weighting space is an innovative way to assess comparability of data sets that could be developed into a new sensitivity analysis tool.

59 BASIC BIOLOGICAL SCIENCES↗

Prediction of vacancy defect diffusion paths in high entropy alloys via machine learning on molecular dynamics data

Identifying the diffusion path of point defects is a critical step in understanding their evolution and the mechanisms of related phenomena. Defect diffusion occurs at small length and time scales, with impacts on material properties that may continue to evolve over ns to μs, ms, and the continuum scale (s, min, etc., and cm, m, etc.). The time scale accessible to molecular dynamics (MD) simulations is limited by small step sizes, typically in the fs range. Thus, surrogate models of MD simulations through machine learning (ML)-based algorithms are of great interest, especially for complex systems such as high entropy alloys (HEAs). In this work, dynamics governing vacancy migration in HEA were approximated with graph convolutional network (GCN) models as ansatzes for kinetic Monte Carlo (KMC) rate catalogs. Network design considered that diffusion in crystalline solids generally depends on interactions between defects and their immediate neighbor atoms. Graphs represented the vacancy surroundings, MD-generated trajectories provided training and comparison datasets, and unsupervised GCN models approximated interatomic dynamics governing vacancy migration in HEAs as ansatzes for KMC. A proof-of-concept model trained on MD data for the Fe, Ni, Cr, Co, and Cu HEA environment was used with two different neighbor interactions to assess the feasibility of training a GCN to predict vacancy defect transition rates in the HEA environment. The resulting setup rapidly generated MD-formatted synthetic trajectories based on dynamics learned from the MD training set, with a time acceleration of roughly two orders of magnitude and a similar diffusion coefficient to MD observations. Additionally, Nudged Elastic Band (NEB) calculations were performed on randomly generated FeNiCrCoCu HEA structures to determine vacancy migration barriers across nearest-neighbor sites. Transition probabilities for each jump, categorized by atomic type, were extracted from these calculations. NEB-based and GCN-based approaches led to similar outcomes.

Reimer, C↗

Integration and Quantitative Comparison of Up-Scaled Molecular Observation Network Data with Existing Soil Databases

MONet provides novel soil molecular data to the research community for understanding biogeochemical processes and complementing other soil datasets. This study assesses MONet's ability to replicate known soil patterns via a comparative analysis of soil respiration (Rs), pH, and clay content against benchmark datasets. Results show moderate agreement for pH and clay content, highlighting MONet's strengths in capturing soil biogeochemical variation in underrepresented regions like urban areas. Rs data are marked by the appropriate trends relative to other datasets, but direct comparison is impractical due to methodological differences in underlying data. Strategic sampling is recommended to improve MONet's coverage and eventual utility in bridging molecular observations with global datasets.

54 ENVIRONMENTAL SCIENCES↗

Clonality, local population structure, and gametophyte sex ratios in cryptic species of the Sphagnum magellanicum complex

Sphagnum (peatmoss) comprises a moss (Bryophyta) clade with approximately 300-500 species. The genus has unparalleled ecological importance because Sphagnum-dominated peatlands store almost a third of the terrestrial carbon pool and peatmosses engineer the formation and microtopography of peatlands. Genomic resources for Sphagnum are being actively expanded, but many aspects of their biology are still poorly known. Among these are the degree to which Sphagnum species reproduce asexually, and the relative frequencies of male and female gametophytes in these haploid-dominant plants. Here, we assess clonality and gametophyte sex ratios and test hypotheses about the local-scale distribution of clones and sexes in four North American species of the S. magellanicum complex. These four species are difficult to distinguish morphologically and are very closely related. We also assess microbial communities associated with Sphagnum host plant clones and sexes at two sites. 405 samples of the four species, representing 57 populations, were subjected to RADseq. Analyses of population structure and clonality based on the molecular data utilized both phylogenetic and phenetic approaches. Multi-locus genotypes (genets) were identified using the RADseq data. Sexes of sampled ramets were determined using a molecular approach that utilized coverage of loci on the sex chromosomes after the method was validated using a sample of plants that expressed sex phenotypically. Sex ratios were estimated for each species, and populations within species. Differences in fitness between genets was estimated as the numbers of ramets each genet comprised. Degree of clonality (numbers of genets/numbers of ramets [samples]) within species, among sites, and between gametophyte sexes were estimated. Sex ratios were estimated for each species, and populations within species. Sphagnum-associated microbial communities were assessed at two sites in relation to Sphagnum clonality and sex. All four species appear to engage in a mixture of sexual and asexual (clonal) reproduction. A single ramet represents most genets but 2-8 ramets were detected for some genets. Only one genet is represented by ramets in multiple populations; all other genets are restricted to a single population. Within populations ramets of individual genets are spatially clustered, suggesting limited dispersal even within peatlands. Sex ratios are male-biased in S. diabolicum but female-biased in the other three species, although significantly so only in S. divinum. Neither species nor males/females differ in levels of clonal propagation. At St. Regis Lake (NY) and Franklin Bog (VT), microbial community composition is strongly differentiated between the sites, but differences between species, genets, and sexes were not detected. Within S. divinum, however, female gametophytes harbored 2-3 times the number oi of microbial taxa as males. These four Sphagnum species all exhibit a similar reproductive patterns that result from a mixture of sexual and asexual reproduction. The spatial patterns of clonally replicated ramets of genets suggest that these species fall between the so-called phalanx patterns where genets abut one another but do not extensively mix, because of limited ramet fragmentation, and the guerrilla patterns where extensive genet fragmentation and dispersal results in greater mixing of different genets. Although sex ratios in bryophytes are most often female-biased, both male and female biases occur in this complex of closely related species. The association of far greater microbial diversity for female gametophytes in S. divinum, which has a female-biased sex ratio, suggests additional research to determine if levels of microbial diversity are consistently correlated with differing patterns of sex ratio biases.

59 BASIC BIOLOGICAL SCIENCES↗

ORNL_AISD_DL-HLgap

This dataset provides supplementary molecular dataset of Deep Learning Workflow for the Inverse Design of Molecules with Specific Optoelectronic Properties. The dataset comprises three main directories such as GDB-9_dataset, Low_HL_Gap_dataset, and High_HL_Gap_dataset which individually has csv files, smiles_txt files, pdb files and xyz files containing information of molecular structures, properties and coordinates generated from deep learning workflow using generative model, surrogate model and DFTB calculation results. GDB-9_dataset contains the molecular data extracted from the original GDB-9 dataset with additional data of DFTB HL gap, surrogate HL gap and molecular property analysis. (the number of atoms, aromaticity and double bond equivalent) Low_HL_Gap_dataset and High_HL_Gap_dataset contains series of dataset for different generations with further split to train and test dataset that were obtained from the iterative workflow described in the manuscript. Additional directory Chemiscope_visualization in Low_HL_Gap_dataset directory contains compressed json files to visualize molecules using chemiscope.org page or application to help readers examine generated molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Biological insights from multi-omics analysis strategies: Complex pleotropic effects associated with autophagy

Research strategies that combine molecular data from multiple levels of genome expression (i.e., multi-omics data), often referred to as a systems biology strategy, has been advocated as a route to discovering gene functions. In this study we conducted an evaluation of this strategy by combining lipidomics, metabolite mass-spectral imaging and transcriptomics data from leaves and roots in response to mutations in two AuTophaGy-related ( ATG ) genes of Arabidopsis . Autophagy is an essential cellular process that degrades and recycles macromolecules and organelles, and this process is blocked in the atg7 and atg9 mutants that were the focus of this study. Specifically, we quantified abundances of ~100 lipids and imaged the cellular locations of ~15 lipid molecular species and the relative abundance of ~26,000 transcripts from leaf and root tissues of WT, atg7 and atg9 mutant plants, grown either in normal (nitrogen-replete) and autophagy-inducing conditions (nitrogen-deficient). The multi-omics data enabled detailed molecular depiction of the effect of each mutation, and a comprehensive physiological model to explain the consequence of these genetic and environmental changes in autophagy is greatly facilitated by the a priori knowledge of the exact biochemical function of the ATG7 and ATG9 proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Selectivity of enzymes involved in the formation of opposite enantiomeric series of p -menthane monoterpenoids in peppermint and Japanese catnip

Peppermint (Mentha x piperita L.) and Japanese catnip (Schizonepeta tenuifolia (Benth.) Briq.) accumulate p-menthane monoterpenoids with identical functionalization patterns but opposite stereochemistry. In the present study, we investigate the enantioselectivity of multiple enzymes involved in monoterpenoid biosynthesis in these species. Based on kinetic assays, mint limonene synthase, limonene 3-hydroxylase, isopiperitenol dehydrogenase, isopiperitenone reductase, and menthone reductase exhibited significant enantioselectivity toward intermediates of the pathway that proceeds through (-)-4S-limonene. Limonene synthase, isopiperitenol dehydrogenase and isopiperitenone reductase of Japanese catnip preferred intermediates of the pathway that involves (+)-4R-limonene, whereas limonene 3-hydroxylase was not enantioselective, and the activities of pulegone reductase and menthone reductase were too low to acquire meaningful kinetic data. Molecular modeling studies with docked ligands generally supported the experimental data obtained with peppermint enzymes, indicating that the preferred enantiomer was aligned well with the requisite cofactor and amino acid residues implicated in catalysis. A striking example for enantioselectivity was peppermint (-)-menthone reductase, which binds (-)-menthone with exquisite affinity but was predicted to bind (+)-menthone in a non-productive orientation that positions its carbonyl functional group at considerable distance to the NADPH cofactor. Here, the work presented here lays the groundwork for structure-function studies aimed at unraveling how enantioselectivity evolved in closely related species of the Lamiaceae and beyond.

59 BASIC BIOLOGICAL SCIENCES↗

Amphiphile Organization in Organic Solutions: An Alternative Explanation for Small-Angle X-ray Scattering Features in Malonamide/Alkane Mixtures

In this work, the role of different intermolecular interactions in the aggregation of amphiphiles in an organic solvent is studied for systems of relevance to liquid-liquid extraction (LLE), a chemical process used to selectively recover metals from complex mixtures. Of specific interest is the role, or lack thereof, of hydrogen bonding, which is often assumed to be a main driver of the organic phase structural organization that has been linked to separation efficacy. Toward that end, a series of malonamide extractants in n-dodecane have been studied in the absence of any extracted aqueous solutes, including water. The series of extractants includes N,N'-dimethyl-N,N'-dibutyltetradecylmalonamide (DMDBTDMA), two of its homologs, and N,N'-dimethyl-N,N'-dioctylhexylethoxymalonamide (DMDOHEMA). This simplified model LLE system enables systematic investigation of the role of dipole-dipole and alkyl tail steric interactions in amphiphile aggregation. Small-angle X-ray scattering (SAXS) profiles computed from molecular dynamics trajectories are in good agreement with experimental SAXS data. Molecular dynamics simulations show that malonamide aggregation results from dipole-driven self-association and lacks characteristic aggregate sizes. Mid-q correlation peaks in the SAXS profiles emerge at high concentration for each malonamide. In those densely packed solutions, the correlation peaks are observed to result from alkyl tail-induced spacing between electron-rich polar head groups, with peak positions determined by the different alkyl tail lengths present in the malonamide molecule. This explanation of the SAXS correlation peaks contrasts with the prevailing literature, which attributes mesoscale features observed in small-angle scattering to the formation of microemulsions. Instead, this work finds that these features are present in the absence of water or any reverse micellar organization of the malonamides. As such, molecular-scale malonamide self-association and packing, rather than microemulsion-based colloidal-scale descriptions, is a more appropriate framework for these LLE systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Water chemistry in flume channel and hyporheic zone (i.e., porewater) associated with: “Rethinking Aerobic Respiration in the Hyporheic Zone Under Variation in Carbon and Nitrogen Stoichiometry”

Dissolved oxygen (DO), total organic carbon (TOC), total nitrogen (TN), molecular data for organic matter, and biochemical reactions for surface water and porewater (i.e., hyporheic zone) collected from a water recirculating flume located at the University of Texas, Austin. The flume contained real river water from Lower Colorado River(Austin, TX) and clean sand. Hyporheic exchange in the flume was induced through The study aims to understand relationships between aerobic metabolism of organic matter and molecular characteristics of organic matter, such as thermodynamic signature and nitrogen content, through the extent of the hyporheic zone at 10 cm- resolution, and through time. During the experiment, organic matter (dry leaves) was added to the flume and removed after 24 hours. The water samples were collected before the addition of leaves, at the time of removal of leaves, and at hour 72. The water samples were analyzed using ultrahigh resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) and total organic carbon (TOC) and total nitrogen (TN) analysis. Dissolved oxygen content throughout the surface water and the hyporheic zone of the flume was measured with a large planar optode. This data package is associated with the publication ’ Rethinking Aerobic Respiration in the Hyporheic Zone Under Variation in Carbon and Nitrogen Stoichiometry’ published in Environmental Science and Technology (Turețcaia et al., 2023 https://doi.org/10.1021/acs.est.3c04765). The dataset is comprised of five folders (1) Diss_O2_pic, (2) input_files (3) output_files; (4) python_code; and (5) R_code . Diss_O2_pic contains siximages of dissolved oxygen distribution in a bedform at hours 0, 24, and 72 of the experiment conducted in a large recirculation flume. Images are in separate R and G channels (i.e., RGB). The input_files contains (1) a csv file with FTICR peaks identified within each sample, (2) a csv file with molecular information pertinent to FTICR data with Gibbs free energy calculations adjusted for environmental temperature, (3) a csv file containing concentrations of non-purgeable organic carbon measured throughout the experiment , (4) a csv file containing concentrations of total nitrogen measured throughout the experiment, (5) a csv file containing total biochemical reactions (i.e., transformations) identified in the dataset, (6) a csv containing transformation profiles, and (7) a csv file containing transformations with formulas, and (8) a jpg file with schematic representation of locations for sample collection. The output_files contains (1) and xlsx file containing percent biochemical reactions containing nitrogen identified across all 39 sample, (2) a csv file of merged FTICR data and molecular information files, (3) a csv files containing average Gibbs free energy within sampling domains and at each sampling location, (4) a csv file with average concentrations of dissolved oxygen across sampling locations at hour 0, (5) a csv file with average concentrations of dissolved oxygen across sampling locations at hour 24, (6) a csv file with average concentrations of dissolved oxygen across sampling locations at hour 72, (7) a csv file with percent chemical classes identified across sampling locations at hour 0, (8) a csv file with percent chemical classes identified across sampling locations at hour 24, (9) a csv file with percent chemical classes identified across sampling locations at hour 72, and (10) a csv file containing percent nitrogen containing biochemical reactions identified across sampling locations at hours 0, 24, and 72. The python_code contains seven ipynb files which are Jupyter Notebooks used for data analysis and figures generation. The R_code contains 3 R files with R code used for data analysis and figures generation. This data package contains the processed data used in the associated manuscript. This data has not been previously published.

54 ENVIRONMENTAL SCIENCES↗