Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reference call set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Structural variant analysis of a cancer reference cell line sample using multiple sequencing technologies

The cancer genome is commonly altered with thousands of structural rearrangements including insertions, deletions, translocation, inversions, duplications, and copy number variations. Thus, structural variant (SV) characterization plays a paramount role in cancer target identification, oncology diagnostics, and personalized medicine. As part of the SEQC2 Consortium effort, the present study established and evaluated a consensus SV call set using a breast cancer reference cell line and matched normal control derived from the same donor, which were used in our companion benchmarking studies as reference samples. We systematically investigated somatic SVs in the reference cancer cell line by comparing to a matched normal cell line using multiple NGS platforms including Illumina short-read, 10X Genomics linked reads, PacBio long reads, Oxford Nanopore long reads, and high-throughput chromosome conformation capture (Hi-C). We established a consensus SV call set of a total of 1788 SVs including 717 deletions, 230 duplications, 551 insertions, 133 inversions, 146 translocations, and 11 breakends for the reference cancer cell line. To independently evaluate and cross-validate the accuracy of our consensus SV call set, we used orthogonal methods including PCR-based validation, Affymetrix arrays, Bionano optical mapping, and identification of fusion genes detected from RNA-seq. We evaluated the strengths and weaknesses of each NGS technology for SV determination, and our findings provide an actionable guide to improve cancer genome SV detection sensitivity and accuracy. A high-confidence consensus SV call set was established for the reference cancer cell line. A large subset of the variants identified was validated by multiple orthogonal methods.

59 BASIC BIOLOGICAL SCIENCES↗

Similarity Downselection: Finding the n Most Dissimilar Molecular Conformers for Reference-Free Metabolomics

Computational methods for creating in silico libraries of molecular descriptors (e.g., collision cross sections) are becoming increasingly prevalent due to the limited number of authentic reference materials available for traditional library building. These so-called “reference-free metabolomics” methods require sampling sets of molecular conformers in order to produce high accuracy property predictions. Due to the computational cost of the subsequent calculations for each conformer, there is a need to sample the most relevant subset and avoid repeating calculations on conformers that are nearly identical. The goal of this study is to introduce a heuristic method of finding the most dissimilar conformers from a larger population in order to help speed up reference-free calculation methods and maintain a high property prediction accuracy. Finding the set of the n items most dissimilar from each other out of a larger population becomes increasingly difficult and computationally expensive as either n or the population size grows large. Because there exists a pairwise relationship between each item and all other items in the population, finding the set of the n most dissimilar items is different than simply sorting an array of numbers. For instance, if you have a set of the most dissimilar n = 4 items, one or more of the items from n = 4 might not be in the set n = 5. An exact solution would have to search all possible combinations of size n in the population exhaustively. We present an open-source software called similarity downselection (SDS), written in Python and freely available on GitHub. SDS implements a heuristic algorithm for quickly finding the approximate set(s) of the n most dissimilar items. We benchmark SDS against a Monte Carlo method, which attempts to find the exact solution through repeated random sampling. We show that for SDS to find the set of n most dissimilar conformers, our method is not only orders of magnitude faster, but it is also more accurate than running Monte Carlo for 1,000,000 iterations, each searching for set sizes n = 3–7 out of a population of 50,000. We also benchmark SDS against the exact solution for example small populations, showing that SDS produces a solution close to the exact solution in these instances. Using theoretical approaches, we also demonstrate the constraints of the greedy algorithm and its efficacy as a ratio to the exact solution.

97 MATHEMATICS AND COMPUTING↗

DMTN-277: The Monster: A reference catalog with synthetic ugrizy-band fluxes for the Vera C. Rubin observatory

In order to facilitate bootstrap photometric calibrations of early Rubin Observatory data we have created an all sky reference catalog called The Monster. This reference catalog uses a rank-ordered set of other reference catalogs to generate synthetic ugrizy-band fluxes that can be used calibrate images processed with the LSST science pipelines. This document describes the methodology used to create The Monster, documents the input external reference catalogs, and performs basic data validation of the first version of The Monster.

79 ASTRONOMY AND ASTROPHYSICS↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

Path-Integrated X-Ray Digital Image Correlation using Synthetic Reference Images

X-rays can provide images when an object is visibly obstructed, allowing for motion measurements via x-ray digital image correlation (DIC). However, x-ray images are path-integrated and contain data for all objects between the source and detector. If multiple objects are present in the x-ray path, conventional DIC algorithms may fail to correlate the x-ray images. A new DIC algorithm called path-integrated (PI)-DIC addresses this issue by reformulating the matching criterion for DIC to account for multiple, independently-moving objects. PI-DIC requires a set of reference x-ray images of each independent object. However, due to experimental constraints, such reference images might not be obtainable from the experiment. Here, this work focuses on the reliability of synthetically-generated reference images, in such cases. A simplified exemplar is used for demonstration purposes, consisting of two aluminum plates with tantalum x-ray DIC patterns undergoing independent rigid translations. Synthetic reference images based on the “as-designed” DIC patterns were generated. However, PI-DIC with the synthetic images suffered some biases due to manufacturing defects of the patterns. A systematic study of seven identified defect types found that an incorrect feature diameter was the most influential defect. Synthetic images were re-generated with the corrected feature diameter, and PI-DIC errors were improved by a factor of 3-4. Final biases ranged from 0.00-0.04 px, and standard uncertainties ranged from 0.06-0.11 px. In conclusion, PI-DIC accurately measured the independent displacement of two plates from a single series of path-integrated x-ray images using synthetically-generated reference images, and the methods and conclusions derived here can be extended to more generalized cases involving stereo PI-DIC for arbitrary specimen geometry and motion. This work thus extends the application space of x-ray imaging for full-field DIC measurements of multiple surfaces or objects in extreme environments where optical DIC is not possible.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Integration of Utility Distributed Energy Resource Management System and Aggregators for Evolving Distribution System Operators

With the rapid integration of distributed energy resources (DERs), distribution utilities are faced with new and unprecedented issues. New challenges introduced by high penetration of DERs range from poor observability to overload and reverse power flow problems, under-over-voltages, maloperation of legacy protection systems, and requirements for new planning procedures. Distribution utility personnel are not adequately trained, and legacy control centers are not properly equipped to cope with these issues. Fortunately, distribution energy resource management systems (DERMSs) are emerging software technologies aimed to provide distribution system operators (DSOs) with a specialized set of tools to enable them to overcome the issues caused by DERs and to maximize the benefits of the presence of high penetration of these novel resources. However, as DERMS technology is still emerging, its definition is vague and can refer to very different levels of software hierarchies, spanning from decentralized virtual power plants to DER aggregators and fully centralized enterprise systems (called utility DERMS). Although they are all frequently simply called DERMS, these software technologies have different sets of tools and aim to provide different services to different stakeholders. This paper explores how these different software technologies can complement each other, and how they can provide significant benefits to DSOs in enabling them to successfully manage evolving distribution networks with high penetration of DERs when they are integrated together into the control centers of distribution utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Human Proteoform Project: Defining the human proteome

Proteins are the primary effectors of function in biology, and thus, complete knowledge of their structure and properties is fundamental to deciphering function in basic and translational research. The chemical diversity of proteins is expressed in their many proteoforms, which result from combinations of genetic polymorphisms, RNA splice variants, and posttranslational modifications. This knowledge is foundational for the biological complexes and networks that control biology yet remains largely unknown. We propose here an ambitious initiative to define the human proteome, that is, to generate a definitive reference set of the proteoforms produced from the genome. Several examples of the power and importance of proteoform-level knowledge in disease-based research are presented along with a call for improved technologies in a two-pronged strategy to the Human Proteoform Project.

59 BASIC BIOLOGICAL SCIENCES↗

$\mathrm{RADICAL}$-Pilot and $\mathrm{PMIx}$/$\mathrm{PRRTE}$: Executing Heterogeneous Workloads at Large Scale on Partitioned $\mathrm{HPC}$ Resources

Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.

97 MATHEMATICS AND COMPUTING↗

Populus_trichocarpa_Breeding_Population_SNPs

These data are from the manuscript “Application of Genomic Prediction in a Populus trichocarpa Breeding Program”, by Brian J. Stanton, David Macaya-Sanz, Chanaka Roshan Abeyratne, David Kainer, Kathy Haiby, Austin Himes, Carlos Gantz, Gerald A. Tuskan, and Stephen P. DiFazio. The data are based on genome resequencing to approximately 10X depth on two collections of Populus trichocarpa trees from Oregon, Washington, California, and British Columbia. The first collection consists of 293 genets collected by Poplar Innovations LLC for a breeding program. The second collection consists of 961 trees collected for the purpose of genome-wide association studies. These genets were sequenced using short, paired-end Illumina sequence reads (Chhetri et al. 2019). Reads were aligned to the P. trichocarpa ′Stettler-14′ reference (Hofmeister et al. 2020), with minor modifications to correct mis-assemblies (Zhou et al. 2020), and variants were called as per methods described in (Abeyratne et al. 2023). Identified variants were filtered using GATK’s VariantFiltration tool (DePristo et al. 2011), with filter expression flag set to “AF < 0.01 || AF > 0.99 || QD < 10.0 || ExcessHet > 20.0 || FS > 10.0 || MQ < 58.0”. SNPs with severe departures from Hardy−Weinberg expectations (exact-test p< 0.01) were also removed using vcftools --hwe flag (Danecek et al. 2011), resulting in 15,627,211 bi-allelic SNPs. The data included here consist of 141,903 high quality bi-allelic genome-wide SNPs obtained by further filtering the original SNP dataset using vcftools with flags --maf 0.05, --max-maf 0.95, --max-missing 0.95, --min-meanDP 10.75, --max-meanDP 43.00, --thin 2000. Collectively, these filtering parameters removed SNPs with 1) a minor allele frequency ≤ 0.05; 2) proportion of missing data for individual loci exceeding 5%; 3) sequencing depth more than 2X mean-depth or less than 0.5X mean-depth; or 4) a distance of

09 BIOMASS FUELS↗

Stability analysis of a one-dimensional multiphysics model of a molten salt fast reactor

Reactor designs with a fast neutron spectrum have recently moved into the research focus for possible Generation IV reactor concepts - a representative of which is the molten salt fast reactor (MSFR). In order to assess the safety and reliability of this reactor concept, it is necessary to obtain knowledge of the system behaviour. An important part of this research concerns the dynamical stability of the MSFR. Research regarding the MSFR stability so far has either been using linearized models for stability criteria or only been examining certain transients at fixed parameter values. This work delivers a comprehensive stability analysis for a non-linearized MSFR model, and a wide range of values for all relevant parameters. For this purpose, a one-dimensional MSFR model was set up, taking neutron kinetics and thermal hydraulics in the form of a system of differential equations into account. The stability of this model was investigated by means of the numerical tool MATCONT, which was used to monitor the evolution of a so-called fixed-point solution, here referring to the steady state at operating conditions of the system, while system parameters got varied. The MATCONT analyses showed no loss of stability for any of the considered parameter variations and no solutions that might co-exist in parallel with the steady-state fixed-point solution. Therefore these results indicate a stable fixed point to which all solution transients converge, making the considered MSFR model insusceptible to deviations from the equilibrium. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

SIMD Programming for the SMASH Shock Physics Code

Many modern CPUs that are available to the NNSA as mission computing resources support vector instruction sets. Making good use of vector instructions, referred to as “vectorization”, is often critical to getting the best performance from these CPUs. While other codes choose to rely on compiler auto-vectorization, the SMASH shock physics code chooses to leverage APIs for explicit vectorization. These APIs are similar to directly calling the CPU vendor’s vector intrinsics, with the additional benefit of being vendor-agnostic. This document explains what the SIMD APIs are and how to use them in developing SMASH.

97 MATHEMATICS AND COMPUTING↗

ARC disruption physics and strategy

Commonwealth Fusion Systems (CFS) plans to operate a tokamak power plant called ARC in the early 2030s. Tokamak plasmas have stability limits that, if crossed, lead to a rapid termination of the plasma, referred to as a disruption. Disruptions pose a melt risk to the first wall resulting from thermal and non-thermal particle heat fluxes, and an electromagnetic loading risk on all metal components within the equilibrium coils. A comprehensive set of models is used herein to provide an assessment of both mitigated and unmitigated ARC disruption loads. A preliminary massive gas injection system is baselined and a runaway electron mitigation coil option is proposed to close possible gaps in the baseline. It is predicted that all ARC disruption loads are within a factor of 2 of the disruption loads in SPARC, a tokamak presently under construction by CFS, and therefore SPARC provides an opportunity to calibrate models, test solutions and inform the design of ARC. The goal for ARC is disruption-free operation, however, the pragmatic design target is to withstand one mitigated disruption per day, and to restart the plasma following mitigation in tens of seconds without interrupting the power output. Unmitigated disruptions must be rare, and experience with unmitigated disruption impacts in SPARC will better define what rare means. The implications of this strategy for plasma disruptivity and disruption prediction are discussed, and operating the ARC scenario on SPARC is expected to refine the ARC final design and operational plan.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

AlloSHP: deconvoluting single homeologous polymorphism for phylogenetic analysis of allopolyploids

Background The genomic and evolutionary study of allopolyploid organisms involves multiple copies of homeologous chromosomes, making their assembly, annotation, and phylogenetic analysis challenging. Bioinformatics tools and protocols have been developed to study polyploid genomes, but sometimes require the assembly of their genomes, or at least the genes, limiting their use. Results We have developed AlloSHP, a command-line tool for detecting and extracting single homeologous polymorphisms (SHPs) from the subgenomes of allopolyploid species. This tool integrates three main algorithms, WGA, VCF2ALIGNMENT and VCF2SYNTENY, and allows the detection of SHPs for the study of diploid-polyploid complexes with available diploid progenitor genomes, without assembling and annotating the genomes of the allopolyploids under study. AlloSHP has been validated on three diploid-polyploid plant complexes, Brachypodium, Brassica, and Triticum-Aegilops, and a set of synthetic hybrid yeasts and their progenitors of the genus Saccharomyces. The results and congruent phylogenies obtained from the four datasets demonstrate the potential of AlloSHP for the evolutionary analysis of allopolyploids with a wide range of ploidy and genome sizes. Conclusions AlloSHP combines the strategies of simultaneous mapping against multiple reference genomes and syntenic alignment of these genomes to call SHPs, using as input data a single VCF file and the reference genomes of the known or closest extant diploid progenitor species. This novel approach provides a valuable tool for the evolutionary study of allopolyploid species, both at the interspecific and intraspecific levels, allowing the simultaneous analysis of a large number of accessions and avoiding the complex process of assembling polyploid genomes.

Allopolyploids↗

Application of FARM to an IES scenario within the FORCE ecosystem

The FARM (Feasible Actuator Range Modifier) software module is a component of the RAVEN-based FORCE framework for analysis of Integrated Energy Systems (IES). FARM supports the HERON software module in the evaluation of the optimal dispatch by evaluating feasible set-points for the different IES unit components. Set-points are required to satisfy limits on both production variables (i.e., the variables to be optimized such as the electrical power, the hydrogen production rate, etc.) and process variables tied to the service life of equipment (e.g., steam flowrate, vessel pressure, turbine firing temperature, etc.). This problem is addressed by adopting a two-stage approach. First, the HERON power dispatcher determines set-points that meet the constraints on the former variables (e.g., power levels and power ramp rate limits). These constraints are called explicit constraints. Then, FARM adjusts these set-points to ensure the respect of the limits on the latter variables given the knowledge of the system physics acquired through machine learning algorithms. These constraints are called implicit constraints. From this standpoint, FARM constitutes a bridge between the HERON power dispatcher that adopts a simplified description of the IES unit (low-resolution physics) and the HYBRID high-fidelity models (high-resolution physics). The original version of the FARM software module (FARM-Alpha) was released by Argonne National Laboratory in January 2021. In the latest version of the code released in July 2022 (FARM-Delta), the Reference Governor (RG) algorithm was upgraded to a Multi-Input Multi-Output version from its original Single-Input Single-Output form. The RG algorithm acts to enforce constraints. With this improvement an IES unit is now treated as a single dynamic system from the standpoint of control. The crosstalk among components in an IES unit is now fully considered thereby ensuring a true optimization is obtained for those units that have multiple set-points. In this report, the capabilities of FARM-Delta operating within the FORCE ecosystem are demonstrated for an IES test case. The specific configuration of IES unit for this case was selected by the IES team with consultation from the Advanced Reactor IES Expert Group. A full TEA analysis that invoked HERON, HYBRID, FARM, and RAVEN was performed and serves to demonstrate how the latest modification to FARM algorithms (i.e., state variable selection, state-space matrices derivation, set-point verification) can shape setpoints that might otherwise compromise the health of equipment through accelerated wear and tear. In this specific test case, it was demonstrated that these algorithms ensure a more efficient utilization of steam resources to be shared by two different subsystems, namely Balance of Plant (BOP) and High-Temperature Steam Electrolysis (HTSE). Finally, some code improvements that can further enhance the user-friendliness are suggested.

97 MATHEMATICS AND COMPUTING↗

Towards a machine-readable literature: finding relevant papers based on an uploaded powder diffraction pattern

A prototype application for machine-readable literature is investigated. The program is called pyDataRecognition and serves as an example of a data-driven literature search, where the literature search query is an experimental data set provided by the user. The user uploads a powder pattern together with the radiation wavelength. The program compares the user data to a database of existing powder patterns associated with published papers and produces a rank ordered according to their similarity score. The program returns the digital object identifier and full reference of top-ranked papers together with a stack plot of the user data alongside the top-five database entries. The paper describes the approach and explores successes and challenges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Transferability of data-driven, many-body models for CO 2 simulations in the vapor and liquid phases

Here, extending on the previous work by Riera et al. [J. Chem. Theory Comput. 16, 2246–2257 (2020)], we introduce a second generation family of data-driven many-body MB-nrg models for CO 2 and systematically assess how the strength and anisotropy of the CO 2 –CO 2 interactions affect the models’ ability to predict vapor, liquid, and vapor–liquid equilibrium properties. Building upon the many-body expansion formalism, we construct a series of MB-nrg models by fitting one-body and two-body reference energies calculated at the coupled cluster level of theory for large monomer and dimer training sets. Advancing from the first generation models, we employ the charge model 5 scheme to determine the atomic charges and systematically scale the two-body energies to obtain more accurate descriptions of vapor, liquid, and vapor–liquid equilibrium properties. Challenges in model construction arise due to the anisotropic nature and small magnitude of the interaction energies in CO 2 , calling for the necessity of highly accurate descriptions of the multidimensional energy landscape of liquid CO 2 . These findings emphasize the key role played by the training set quality in the development of transferable, data-driven models, which, accurately representing high-dimensional many-body effects, can enable predictive computer simulations of molecular fluids across the entire phase diagram.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

CORPSE model with litter decomposition parameters derived from the LIDET dataset

This is a version of the CORPSE model (Carbon, Organisms, Rhizosphere and Protection in the Soil Environment, Sulman et al. 2014) that uses litter decomposition parameters derived from a modified Monte Carlo simulation using the LIDET litter decomposition dataset (Long-term Intersite Decomposition Experiment Team, Harmon 2013). The code also includes the Baseline parameters, and the eight other best parameter sets identified in a modified Monte Carlo simulation. Related publication:Juice, S.M., Ridgeway, J.R., Hartman, M.D., Parton, W.J., Berardi, D.M., Sulman, B.N., Allen, K.E., & Brzostek, E.R. Reparameterizing litter decomposition using a simplified Monte Carlo method improves litter decay simulated by a microbial model and alters bioenergy soil carbon estimates. Description of files:The folder "Input Files" contains one folder for each LIDET site with data necessary to run the model. Note that "(site)" in the filenames below indicates where the LIDET site code appears (see Table 1 for site codes). Data streams include: CORPSE_full_spinup_litter.csv, CORPSE_full_spinup_rhizo.csv, CORPSE_full_spinup_bulk.csv, litterbag_init_100g_6spp.csv: initial C and N (kg C or N/m2) pool values for each soil layer, the litterbag_init_100_6spp.csv file is for the litterbag layer and is the same file for all sites. All initial C and N files have the same columns (Column - Description - Units) uFastC - Unprotected fast decomposing carbon - kg carbon/m2 uSlowC - Unprotected slow decomposing carbon - kg carbon/m2 uNecroC - Unprotected necromass carbon - kg carbon/m2 pFastC - Protected fast decomposing carbon - kg carbon/m2 pSlowC - Protected slow decomposing carbon - kg carbon/m2 pNecroC - Protected necromass carbon - kg carbon/m2 livingMicrobeC - Carbon in living microbial biomass - kg carbon/m2 uFastN - Unprotected fast decomposing nitrogen - kg nitrogen/m2 uSlowN - Unprotected slow decomposing nitrogen - kg nitrogen/m2 uNecroN - Unprotected necromass nitrogen - kg nitrogen/m2 pFastN - Protected fast decomposing nitrogen - kg nitrogen/m2 pSlowN - Protected slow decomposing nitrogen - kg nitrogen/m2 pNecroN - Protected necromass nitrogen - kg nitrogen/m2 inorganicN - Inorganic nitrogen - kg nitrogen/m2 CO2 - Carbon in carbon dioxide - kg carbon/m2 livingMicrobeN - Nitrogen in living microbial biomass - kg nitrogen/m2 soilT (site) DOY274start.csv: Average daily soil temperature (oC) interpolated from previously calculated monthly values used in DayCent LIDET simulations (Bonan et al., 2013). soilT (site) DOY274start.csv: Average daily soil volumetric water content (VWC) scalar interpolated from previously calculated monthly values used in DayCent LIDET simulations (Bonan et al., 2013). litter production.csv: Average daily litter production values for each site, data sources listed in Table S3 of related publication. litter (site) CN.csv: C:N ratio for each species from LIDET dataset (Table 2, Harmon 2013). (site).csv: Table indicating number of observations for each species decomposed at each site. Instructions: Save the model code ("CORPSE_LIDET.R") and "Input Files" folder in the same folder. Also make a folder for the model output (e.g., "results_Baseline") in the same folder. Set the working directory (setwd) in the model code to the folder with the files saved in step #1. Select the parameter set to use for the litter and litterbag compartments, comment out all other parameter sets. Run code. Output will be saved in the folder made in step 1. Output destination can be changed as necessary in code section called "Running the model." Table 1 LIDET sites and site codes used in model files. Site Code - Site AND - H.J. Andrews Experimental Forest BNZ - Bonanza Creek Experimental Forest BSF - Blodgett Research Forest CDR - Cedar Creek Natural History Area CPR - Central Plains Experimental Range HBR - Hubbard Brook Experimental Forest HFR - Harvard Forest JUN - Juneau KBS - Kellogg Biological Station KNZ - Konza Prairie Research Natural Area NWT - Niwot Ridge/Green Lakes Valley OLY - Olympic National Park OLY Conifer forest SEV - Sevilleta National Wildlife Refuge SMR - Santa Margarita Ecological Reserve UFL - University of Florida VCR - Virginia Coast Reserve Table 2 LIDET species and species codes used in model files (6 common species). Species - Species Code Sugar maple (Acer saccharum) - ACSA Drypetes (Drypetes glauca) - DRGL Red pine (Pinus resinosa) - PIRE Chestnut oak (Quercus prinus) - QUPR Western redcedar (Thuja plicata) - THPL Wheat (Triticum aestivum) - TRAE References:Bonan, G. B., Hartman, M. D., Parton, W. J., & Wieder, W. R. (2013). Evaluating litter decomposition in earth system models with long-term litterbag experiments: an example using the Community Land Model version 4 (CLM4). Global Change Biology, 19(3), 957-974. https://doi.org/https://doi.org/10.1111/gcb.12031 Harmon, M. (2013). LTER Intersite Fine Litter Decomposition Experiment (LIDET), 1990 to 2002. Long-Term Ecological Research. Forest Science Data Bank, Corvallis, OR. [Data set]. Accessed http://andlter.forestry.oregonstate.edu/data/abstract.aspx?dbcode=TD023. https://doi.org/10.6073/pasta/f35f56bea52d78b6a1ecf1952b4889c5. Sulman, B. N., Phillips, R. P., Oishi, A. C., Shevliakova, E., & Pacala, S. W. (2014). Microbe-driven turnover offsets mineral-mediated storage of soil carbon under elevated CO2. Nature Climate Change, 4, 1099 - 1102. https://doi.org/10.1038/nclimate2436

Juice, Stephanie↗