Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

The final WaZP galaxy cluster catalog of the Dark Energy Survey and comparison with SZE data

In this work, we present and characterize the galaxy cluster catalog detected by the WaZP cluster finder, which is not based on red-sequence identification, on the full six years of observations of the Dark Energy Survey (DES-Y6). The full catalog contains over 400k detected clusters with richnesses, Ngals, above 5 and that reach redshifts up to 1.3. We also provide a version of the catalog where the observation depth and richness computation are homogenized to be used for cosmology, containing 33k rich (Ngals >25) clusters. We compare our results with the previous WaZP catalog obtained from the DES first-year data release (DES-Y1). We find that essentially all clusters within the common footprint and depth limit are recovered. The deeper observations on DES-Y6 and the more complete available spectroscopic redshift sample lead to improvements in the redshifts of the clusters, resulting in an average scatter of 1.4% and offset of 0.2%. The optical clusters are also cross-matched with Sunyaev Zel'dovich Effect (SZE) cluster samples detected by the South Pole Telescope (SPT) and the Atacama Cosmology Telescope (ACT). We find that essentially all SZE clusters with reasonable overlapping footprint have a corresponding WaZP cluster. Conversely, 90% of the optical detections with richness greater than 150 have a counterpart in the deeper regions of the SZE surveys. Based on cross-match with the SZE catalogs, we also find that 15-20% of the SZE matched systems have more than one possible WaZP counterpart at the same redshift and within the SZE R500c, indicating possible interacting or unrelaxed systems. Finally, given the optical and SZE beams, WaZP and SZE centerings are found to be consistent. A more detailed study of the SZE-WaZP mass-richness relation will be presented in a separate paper.

Benoist, C. [OCA, Nice, Lab. Lagrange; LIneA, Rio ↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

Hardware In the Loop for Demand Flexibility (HIL4DF) v1.0

The software package in question is a collection of simulation models in the Modelica language, representing a variety of mechanical system designs and envelope conditions related to LBL's FLEXLAB facility. The collection of models also features multiple controls sequences that can be simulated with the FLEXLAB model to simulate different demand flexibility scenarios. Additionally, this package will feature datasets from 3 experimental tests, used for calibration, validation and comparison against the Modelica models, this includes weather data that can be used to replicate different scenarios in simulation across the same weather conditions experienced in real experiments. Given FLEXLAB high level of instrumentation and available data, the models are calibrated across multiple measurement points, and thus results from the extension of this model to other climate zones or control sequences, would provide high level of confidence.

Huang, Weiping↗

Nuclear Data Libraries Sensitivity Studies for ORSA Using SCALE

Subcritical assemblies offer valuable training capabilities in nuclear criticality safety (NCS) for individuals handling fissile material. At Oak Ridge National Laboratory(ORNL), the Oak Ridge Subcritical Assembly (ORSA), a new experimental facility, is being established to provide hands-on training for the Nuclear Criticality Safety Program (NCSP).It is essential to accurately determine the neutron multiplication factor (keff) to ensure that ORSA remains subcritical and safe during operations. This study investigated the sensitivity of keff to variations across Evaluated Nuclear Data File (ENDF/B) libraries, consisting of ENDF/B-VII.1, ENDF/B-VIII.0, and ENDF/B-VIII.1. The analysis was conducted using the CSAS6 sequence in the SCALE-6.3 code system. Individual isotopes in the ORSA model were replaced one at a time with ENDF/B-VII.1 as the base library and changing to ENDF/B-VIII.0or ENDF/B-VIII.1. The results demonstrated that the changes in keffof the nuclides associated with the ORSA model were mostly within the uncertainty of the base model (~24 pcm), except for primary nuclides like Uranium-235and H-poly with few other nuclides. The relative delta keff values of Uranium-235 and H-poly, expressed in pcm, were +280 and -212 in ENDF/B-VIII.0 and +332 and -275 in ENDF/B-VIII.1, respectively, which were notable changes in reactivity. These results demonstrated that ORSA was largely insensitive to variations across these nuclear data libraries.

Hong, Evan [North Carolina State University]↗

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

Salt supplementation-induced metabolic reprogramming in Streptomyces coelicolor

Members of the genus Streptomyces are major producers of a wide variety of secondary metabolites that serve as bioactive compounds. Many secondary metabolites are produced in response to environmental signals such as biotic and abiotic stresses. In this study, we identified salt supplementation as one of the stimuli activating secondary metabolism in the model Streptomyces species, Streptomyces coelicolor. Comparative metabolomics revealed overproduction of several known secondary metabolites, most notably undecylprodigiosin and coelimycin P1, in addition to their biosynthetic intermediates and derivatives, as well as many unknown metabolites. Transcriptomic analysis revealed activation of diverse biological processes including cation uptake, compatible solute production, and the phosphate limitation stress response through conserved and species-specific mechanisms, presumably to overcome the increased salinity. This response leads to activation of a variety of regulatory and metabolic pathways required for production of secondary metabolites including activation of conserved metabolic pathways for energy and substrate supply and species-specific secondary metabolite biosynthetic gene clusters. Furthermore, several promoter sequences contributing to upregulation of secondary metabolism induced by salt supplementation were identified. Overall, our data show how S. coelicolor copes with the increased salinity and tailors the cellular metabolism toward secondary metabolism in a conserved and species-specific manner.

Otani, Hiroshi [USDOE Joint Genome Institute (JGI)↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

Towards verifiable cancer digital twins: tissue level modeling protocol for precision medicine

Cancer exhibits substantial heterogeneity, manifesting as distinct morphological and molecular variations across tumors, which frequently undermines the efficacy of conventional oncological treatments. Developments in multiomics and sequencing technologies have paved the way for unraveling this heterogeneity. Nevertheless, the complexity of the data gathered from these methods cannot be fully interpreted through multimodal data analysis alone. Mathematical modeling plays a crucial role in delineating the underlying mechanisms to explain sources of heterogeneity using patient-specific data. Intra-tumoral diversity necessitates the development of precision oncology therapies utilizing multiphysics, multiscale mathematical models for cancer. This review discusses recent advancements in computational methodologies for precision oncology, highlighting the potential of cancer digital twins to enhance patient-specific decision-making in clinical settings. We review computational efforts in building patient-informed cellular and tissue-level models for cancer and propose a computational framework that utilizes agent-based modeling as an effective conduit to integrate cancer systems models that encode signaling at the cellular scale with digital twin models that predict tissue-level response in a tumor microenvironment customized to patient information. Furthermore, we discuss machine learning approaches to building surrogates for these complex mathematical models. These surrogates can potentially be used to conduct sensitivity analysis, verification, validation, and uncertainty quantification, which is especially important for tumor studies due to their dynamic nature.

60 APPLIED LIFE SCIENCES↗

MATEY: multiscale adaptive transformer models for spatiotemporal physical systems

Accurate representation of the multiscale features in spatiotemporal physical systems using vision transformer architectures requires extremely long, computationally prohibitive token sequences. To address this issue, we propose two novel adaptive tokenization schemes that dynamically adjust patch sizes based on local features: one ensures convergent behavior to uniform patch refinement, while the other offers better computational efficiency. Moreover, we present a set of spatiotemporal attention schemes, where the temporal or axial spatial dimensions are decoupled, to evaluate their baseline computational and data efficiencies and to determine whether adaptive tokenization can improve this performance. We assess the performance of the proposed multiscale adaptive model, MATEY, in a sequence of experiments. Compared to a full spatiotemporal attention scheme or a scheme that decouples only the temporal dimension, we find that fully decoupled axial attention is less efficient and expressive, requiring more training time and model parameters to achieve the same accuracy. The experiments on the adaptive tokenization schemes show that, compared to a uniformly refined model, the proposed schemes achieve comparable or improved accuracy at a much lower cost in the tested two-dimensional settings. While the asymptotic analysis suggests the potential for favorable scaling, empirical validation at substantially longer sequence lengths remains to be performed in future work. Finally, we demonstrate in two fine-tuning tasks featuring different physics that models pretrained on PDEBench data outperform the ones trained from scratch, especially in the low data regime with frozen attention.

adaptive tokenization↗

A genomic view of Earth’s biomes

Microorganisms are essential to all life on Earth through critical roles in key biological processes and diverse interactions with other organisms that shape ecosystems, drive biogeochemical cycles and influence both human health and environmental health. High-throughput sequencing from environmental samples has revolutionized the understanding of microbial diversity and functions. With vast amounts of genomes now available across Earth’s biomes, these data provide a blueprint of microbial life that can be harnessed for a more holistic understanding of microbiome structure and function across the various ecosystems on Earth. Here we review the application of genome-centric approaches, including recent advances in single-cell sequencing and functional profiling, to survey microbial and viral diversity. Furthermore, we highlight some of the most impactful evolutionary and functional discoveries, explore the spatial diversity and temporal dynamics of microorganisms across diverse environments, and discuss genome-enabled insights into host-associated microorganisms.

Ecology↗

From subsidies to stressors: Positively skewed ecological gradients alter biological responses to nutrients in streams

Abstract Subsidy–stress gradients offer a useful framework for understanding ecological responses to perturbation and may help inform ecological metrics in highly modified systems. Historic, region‐wide shifts from bottomland hardwood forest to row crop agriculture can cause positively skewed impact gradients in alluvial plain ecoregions, resulting in tolerant organisms that typically exhibit a subsidy response (increased abundance in response to environmental stressors) shifting to a stress response (declining abundance at higher concentrations). As a result, observed biological tolerance in modified ecosystems may differ from less modified regions, creating significant challenges for detecting biological responses to restoration efforts. Using the agriculturally dominated Mississippi Alluvial Plain (MAP) ecoregion in Mississippi, USA, as a case study, we tested the hypothesis that macroinvertebrate taxa that typically display a subsidy response to nutrient enrichment in less modified ecoregions (i.e., nutrient‐tolerance) shift to a stress response to increasing nutrients in highly modified watersheds with elevated baseline nutrient conditions (i.e., nutrient intolerance). The abundance and diversity of MAP‐specific intolerant taxa identified with threshold indicator taxa analysis were either unresponsive or exhibited a subsidy response to increasing nutrients in less modified ecoregions in Mississippi with less land alteration and lower nutrient concentrations, but declined at higher concentrations, providing evidence for a stress response to elevated nutrients in the MAP. Additionally, MAP‐specific tolerant and intolerant taxa richness responded to increased nutrients predictably and consistently across space and time within the MAP. However, in MAP streams, elevated specific conductance was predicted to dampen the response of tolerant and intolerant taxa richness to increasing nutrient concentrations, highlighting the importance of considering multistressor interactions when interpreting biological data. Lastly, we demonstrate the efficacy of this approach with sediment bacterial communities characterized with amplicon sequencing, which lack sufficient life history characteristics necessary for the development of multimetric indices. Both macroinvertebrate and bacterial communities responded similarly to increasing nutrient concentrations, suggesting DNA‐based approaches may provide an efficient biological assessment tool for monitoring water quality improvements in highly modified watersheds.

DeVilbiss, Stephen E. [U.S. Geological Survey Lowe↗

Automated Signal Timing Plan Reconstruction Using High-Resolution Event-Based Controller Data for Digital Twins

Transportation digital twins are essential tools for evaluating emerging technologies such as connected and automated vehicles, adaptive traffic signal control, and mobility optimization strategies. Realistic digital twins require accurate emulation of real-world signal controllers and detailed signal timing plans. However, signal timing plans are often unavailable or difficult to access, forcing researchers and modelers to rely on assumed fixed timings or halt their analysis. To overcome this challenge, we present a method that directly estimates signal timing plan parameters using high-resolution, event-based data from traffic signal controllers. The proposed method extracts key parameters, including cycle length, offset, phase sequence, coordinated phases, phase-specific minimum and maximum green durations, vehicle extensions, and splits under coordination. A rule-based deterministic signal timing reconstruction algorithm based on traffic signal operation rules, such as those outlined in the Signal Timing Manual, is developed and validated. We evaluate this method, which uses high-resolution controller event logs and verified signal timing plans, on 94 signalized intersections in Nashville, Tennessee, demonstrating their ability to generate accurate, simulation-ready signal timing plans for tools such as SUMO and Vissim.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)↗

Using DNA affinity purification sequencing (DAP-seq) to identify in vitro binding sites of potential Novosphingobium aromaticivorans DSM12444 transcription factors

Genome-wide binding sites of 44 putative transcription factors (TFs) from Novosphingobium aromaticivorans DSM12444 were analyzed using DNA affinity purification sequencing. We report that 32 of these TFs have at least one area of enrichment. These data will help better understand aromatic metabolism and other features of N. aromaticivorans biology.

DAP-seq↗

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. RNA structure prediction is not yet possible due to a lack of high-quality reference data associated with organismal phenotypes that could inform RNA function. We present GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB). GARNET links RNA sequences to experimental and predicted optimal growth temperatures of GTDB reference organisms. Using GARNET, we develop sequence- and structure-aware RNA generative models, with overlapping triplet tokenization providing optimal encoding for a GPT-like model. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identify mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB-derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

59 BASIC BIOLOGICAL SCIENCES↗

Populus_trichocarpa_Breeding_Population_SNPs

These data are from the manuscript “Application of Genomic Prediction in a Populus trichocarpa Breeding Program”, by Brian J. Stanton, David Macaya-Sanz, Chanaka Roshan Abeyratne, David Kainer, Kathy Haiby, Austin Himes, Carlos Gantz, Gerald A. Tuskan, and Stephen P. DiFazio. The data are based on genome resequencing to approximately 10X depth on two collections of Populus trichocarpa trees from Oregon, Washington, California, and British Columbia. The first collection consists of 293 genets collected by Poplar Innovations LLC for a breeding program. The second collection consists of 961 trees collected for the purpose of genome-wide association studies. These genets were sequenced using short, paired-end Illumina sequence reads (Chhetri et al. 2019). Reads were aligned to the P. trichocarpa ′Stettler-14′ reference (Hofmeister et al. 2020), with minor modifications to correct mis-assemblies (Zhou et al. 2020), and variants were called as per methods described in (Abeyratne et al. 2023). Identified variants were filtered using GATK’s VariantFiltration tool (DePristo et al. 2011), with filter expression flag set to “AF < 0.01 || AF > 0.99 || QD < 10.0 || ExcessHet > 20.0 || FS > 10.0 || MQ < 58.0”. SNPs with severe departures from Hardy−Weinberg expectations (exact-test p< 0.01) were also removed using vcftools --hwe flag (Danecek et al. 2011), resulting in 15,627,211 bi-allelic SNPs. The data included here consist of 141,903 high quality bi-allelic genome-wide SNPs obtained by further filtering the original SNP dataset using vcftools with flags --maf 0.05, --max-maf 0.95, --max-missing 0.95, --min-meanDP 10.75, --max-meanDP 43.00, --thin 2000. Collectively, these filtering parameters removed SNPs with 1) a minor allele frequency ≤ 0.05; 2) proportion of missing data for individual loci exceeding 5%; 3) sequencing depth more than 2X mean-depth or less than 0.5X mean-depth; or 4) a distance of

09 BIOMASS FUELS↗

Phylogenomics and genetic analysis of solvent-producing Clostridium species

Abstract The genus Clostridium is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 C. beijerinckii , 57 C. saccharobutylicum , 4 C. saccharoperbutylacetonicum , 5 C. butyricum , 7 C. acetobutylicum , and 3 C. tetanomorphum genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or Escherichia coli for enhanced chemical production. Recently, gene sequences from this collection were used to engineer Clostridium autoethanogenum , a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production, as well as butanol, butanoic acid, hexanol and hexanoic acid production.

59 BASIC BIOLOGICAL SCIENCES↗