Search NASA⌕ Search

SEARCH · Search NASA

Results for “genome analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Coupling flux balance analysis with reactive transport modeling through machine learning for rapid and stable simulation of microbial metabolic switching

Integrating genome-scale metabolic networks with reactive transport models (RTMs) provides a detailed description of the dynamic changes in microbial growth and metabolism. Despite promising demonstrations in the past, computational inefficiency has been pointed out as a critical issue to overcome because it requires repeated application of linear programming (LP) to obtain flux balance analysis (FBA) solutions in every time step and spatial grid. To address this challenge, we propose a new simulation method where we train and validate artificial neural networks (ANNs) using randomly sampled FBA solutions and incorporate the resulting surrogate FBA model (represented as algebraic equations) into RTMs as source/sink terms. We demonstrate the efficiency of our method via a case study of Shewanella oneidensis MR-1. During aerobic growth on lactate, S. oneidensis produces metabolic byproducts (such as pyruvate and acetate), which are subsequently consumed as alternative carbon sources when the preferred nutrients are depleted. To effectively simulate these complex dynamics, we used a cybernetic approach that models metabolic switches as the outcome of dynamic competition among multiple growth options. In both zero-dimensional batch and one-dimensional column configurations, the ANN-based surrogate models achieved substantial reduction of computational time by several orders of magnitude compared to the original LP-based FBA models. Moreover, the ANN models produced robust solutions without any special measures to prevent numerical instability. These developments significantly promote our ability to utilize genome-scale networks in complex, multi-physics, and multi-dimensional ecosystem modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive models of the genetic bases underlying budding yeast fitness in multiple environments

Abstract The ability of organisms to adapt and survive depends on the effects of genes and the environment on fitness. However, the multigenic nature of fitness and genotype-by-environment interactions hinder our understanding of the genetic basis of fitness. Here, we established fitness prediction models for 35 environments using machine learning and existing fitness data and different genetic variant types for a Saccharomyces cerevisiae population. Models revealed that the predictive ability of genetic variants varied across environments, with copy number variants explaining the majority of fitness variation in most cases. Model interpretation showed that different variant types identified distinct gene sets associated with predictive variants. These gene sets were significantly enriched in experimentally validated genes affecting fitness in only a subset of environments, indicating that many genes influencing fitness remain unexplored. Notably, non-experimentally validated genes were more important than validated ones for fitness predictions. Gene contributions to predictions were both isolate- and environment-dependent, pointing to gene-by-gene and gene-by-environment interactions. Furthermore, models uncovered experimentally validated and novel candidate genetic interactions for a well-characterized stress, the fungicide benomyl. These findings highlight the feasibility of identifying the genetic basis of fitness by using different genetic variant types and offer novel targets for future functional analysis.

DNA copy number variations↗

Longitudinal genome-wide association study reveals early QTL that predict biomass accumulation under cold stress in sorghum

Sorghum bicolor is a promising cellulosic feedstock crop for bioenergy due to its high biomass yields. However, early growth phases of sorghum are sensitive to cold stress, limiting its planting in temperate environments. Cold adaptability is crucial for cultivating bioenergy and grain sorghum at higher latitudes and elevations, or for extending the growing season. Identifying genes and alleles that enhance biomass accumulation under early cold stress can lead to improved sorghum varieties through breeding or genetic engineering. We conducted image-based phenotyping on 369 accessions from the sorghum Bioenergy Association Panel (BAP) in a controlled environment with early cold treatment. The BAP includes diverse accessions with dense genotyping and varied racial, geographical, and phenotypic backgrounds. Daily, non-destructive imaging allowed temporal analysis of growth-related traits and water use efficiency (WUE). A genome-wide association study (GWAS) was performed to identify genomic intervals and genes associated with cold stress response. The GWAS identified transient quantitative trait loci (QTL) strongly associated with growth-related traits, enabling an exploration of the genetic basis of cold stress response at different developmental stages. This analysis of daily growth traits, rather than endpoint traits, revealed early transient QTL predictive of final phenotypes. The study identified both known and novel candidate genes associated with growth-related traits and temporal responses to cold stress. The identified QTL and candidate genes contribute to understanding the genetic mechanisms underlying sorghum's response to cold stress. These findings can inform breeding and genetic engineering strategies to develop sorghum varieties with improved biomass yields and resilience to cold, facilitating earlier planting, extended growing seasons, and cultivation at higher latitudes and elevations.

59 BASIC BIOLOGICAL SCIENCES↗

Congruence of Clusters Defined By Whole Genome Sequencing and MALDI-TOF for Bacteria Isolated From Cleanrooms

Introduction: Oligotrophic conditions can render cleanrooms inhospitable to microbes. Despite these constraints, fungi and bacteria are frequently isolated from surfaces in astromaterials cleanrooms at the Johnson Space Center. Bacillus species are of particular concern because endospores belonging to this genus are resilient and can affect astromaterials. Current monitoring programs rely on 16S rRNA sequencing and the VITEK2 Compact system. These methods have limited power to resolve Bacillus species. Matrix-assisted laser desorption - time of flight mass spectrometry (MALDI-TOF MS), provides a rapid, low cost, method of identifying bacterial isolates and has a higher resolution than 16S rRNA sequencing, particularly for Bacillus species; however, few studies have compared this method to the industry gold standard, whole genome sequencing (WGS). Methods: Based on 16S rRNA classification, we selected 14 isolates for analysis with MALDI-TOF and WGS. Mass spectra were generated with MALDI-TOF MS and processed with custom scripts to identify clusters of closely related isolates and calculate a matrix of pairwise cosine similarity scores. Hybrid Illumina and Nanopore sequencing were used to generate draft genomes. Pairwise similarity scores were calculated from these genomes based on the average amino acid identity (AAI) predicted from single copy core genes. Congruence of clustering between these methods, was assessed by calculating adjusted Rand and Wallace coefficients. Results: Clusters of species generated from MALDI-TOF MS showed good agreement of phylotypes generated with WGS. Pairs of strains that were > 94% similar to each other in terms of predicted amino acid sequences consistently showed cosine similarities of mass spectra > 0.65 and, of the 9 clusters identified with WGS, 8 were identical with MALDI-TOF. This corresponds to an adjusted Rand index of 0.95 and a 95% confidence interval of 0.80 – 1.00 for adjusted Wallace coefficients. The only discordance was for a pair of isolates that were classified as Paenibacillus species. This pair showed relatively high similarity (0.84) in terms of MALDI-TOF MS but only 85% similarity in terms of AAI. Conclusion: This study shows that MALDI-TOF and WGS exhibit a similar ability to delineate Bacillus species isolated from cleanrooms and taxonomic units described by these two methods are consistent with one another. Since MALDI-TOF MS is low in cost and high in throughput, this approach appears to be an ideal option for routine microbial monitoring and identifying Bacillus species.

Farnaz Mazhari↗

Identifying microbial functional guilds performing cryptic organotrophic and lithotrophic redox cycles in anaerobic granular biofilms

Granular biofilms used in anaerobic digester systems contain diverse microbial populations that interact to hydrolyze organic matter and produce methane within controlled environments. Prior research investigated the feasibility of utilizing granular biofilms obtained from an anaerobic digester to remove nitrate without the addition of exogenous electron donors. These granules possessed a unique structure of alternating light and dark iron sulfide and pyrite rich layers that potentially served as both an electron source and sink, linking carbon, nitrogen, sulfur, and iron cycles. To characterize the functional roles of diverse microbial populations enriched within these layered biofilms, we analyzed metagenomes obtained from three different granules. Comparisons between the functional gene content of forty metagenome assembled genomes (MAGs) identified phylogenetically cohesive functional guilds. Each of these functional MAG clusters was assigned to specific steps in anaerobic digestion (hydrolysis, acidogenesis, acetogenesis, and methanogenesis) and anaerobic respiration (denitrification and sulfate reduction). Comparisons with metagenomes derived from a variety of natural and engineered ecosystems confirmed that the enriched denitrifying bacteria were similar to populations typically found in wetlands and biological nitrogen removal systems. Analysis of read alignments to individual genes within the forty MAGs identified conserved genomic features that were representative of the functions that distinguished functional guilds. Overall, this research illustrates the utility of functional based classification of microorganisms for characterizing ecosystem functions and highlights the potential application of engineered ecosystems to serve as experimental models for complex natural ecosystems.

Ecosystem engineering↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

A haplotype-resolved reference genome for Eucalyptus grandis

Eucalyptus grandis is a hardwood tree used worldwide as pure species or hybrid partner to breed fast-growing plantation forestry crops that serve as feedstocks of timber and lignocellulosic biomass for pulp, paper, biomaterials, and biorefinery products. The current v2.0 genome reference for the species served as the first reference for the genus and has helped drive the development of molecular breeding tools for eucalypts. Using PacBio HiFi long reads and Omni-C proximity ligation sequencing, we produced an improved, haplotype-phased assembly (v4.0) for TAG0014, an early-generation selection of E. grandis. The 2 haplotypes are 571 Mbp (HAP1) and 552 Mbp (HAP2) in size and consist of 37 and 46 contigs scaffolded onto 11 chromosomes (contig N50 of 28.9 and 16.7 Mbp), respectively. These haplotype assemblies are 70-90 Mbp smaller than the diploid v2.0 assembly but capture all except one of the 22 telomeres, suggesting that substantial redundant sequence was included in the previous assembly. A total of 35,929 (HAP1) and 35,583 (HAP2) gene models were annotated, of which 438 and 472 contain long introns (>10 kbp) in gene models previously (v2.0) identified as multiple smaller genes. These and other improvements have increased gene annotation completeness levels from 93.8 to 99.4% in the v4.0 assembly. We found that 6,493 and 6,346 genes are within tandem duplicate arrays (HAP1 and HAP2, respectively, 18.4 and 17.8% of the total) and >43.8% of the haplotype assemblies consists of repeat elements. Analysis of synteny between the haplotypes and the E. grandis v2.0 reference genome revealed extensive regions of collinearity, but also some major rearrangements, and provided a preview of population and pangenome variation in the species.

Lötter, Anneri↗

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS↗

Genomic-based biosurveillance for avian influenza: whole genome sequencing from wild mallards sampled during autumn migration in 2022–2023 reveals a high co-infection rate on migration stopover site in Georgia

The Caucasus region, including Georgia, is an important intersection for migratory waterbirds, offering potential for avian influenza virus (AIV) transmission between populations from different geographic areas. In 2022 and 2023, wild ducks were sampled during autumn migration events in Georgia to study the genetic relationships and molecular characteristics of influenza strains. Sequencing and phylogenetic analysis were used to compare the sampled strains to reference sequences from Africa, Asia, and Europe, allowing assessment of genetic relationships and virus transmission between migratory birds. Protein language modeling identified potential co-infections. Of 225 duck samples, 128 tested positive for the influenza M gene. 55 influenza-positive samples underwent whole-genome sequencing, revealing significant diversity. Analysis of the hemagglutinin (HA) segment showed notable differences among subtypes. Most samples were H6N1 and H6N6, but co-infections with combinations like H6H3, N8N1, N6H9, N2N6, and H9H6/N1N2 were also identified. These findings demonstrate the high variability of influenza viruses in migratory waterbirds in Georgia, including a notable rate of co-infections. Some samples exhibited uncommon genetic characteristics compared to other strains from the same year, suggesting Georgia’s role as a mixing vessel for influenza viruses. This facilitates reassortment during co-infections and contributes to the genetic diversity observed across flyways.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-omics analysis reveals the dynamic interplay between Vero host chromatin structure and function during vaccinia virus infection

The genome folds into complex configurations and structures thought to profoundly impact its function. The intricacies of this dynamic structure-function relationship are not well understood particularly in the context of viral infection. To unravel this interplay, here we provide a comprehensive investigation of simultaneous host chromatin structural (via Hi-C and ATAC-seq) and functional changes (via RNA-seq) in response to vaccinia virus infection. Over time, infection significantly impacts global and local chromatin structure by increasing long-range intra-chromosomal interactions and B compartmentalization and by decreasing chromatin accessibility and inter-chromosomal interactions. Local accessibility changes are independent of broad-scale chromatin compartment exchange (~12% of the genome), underscoring potential independent mechanisms for global and local chromatin reorganization. While infection structurally condenses the host genome, there is nearly equal bidirectional differential gene expression. Despite global weakening of intra-TAD interactions, functional changes including downregulated immunity genes are associated with alterations in local accessibility and loop domain restructuring. Therefore, chromatin accessibility and local structure profiling provide impactful predictions for host responses and may improve development of efficacious anti-viral counter measures including the optimization of vaccine design.

59 BASIC BIOLOGICAL SCIENCES↗

Role of Tensile Stress in DNA Nanoresonators for Epigenetic Studies

The evaluation of epigenetic features such as DNA methylation is becoming increasingly important in many biochemical processes like gene expression and transcription as well as in several diseases like schizophrenia or diabetes. Here, in this work, we report that self-assembled nanomechanical resonators entirely composed of DNA molecules can be used to explore gross changes in DNA methylation levels (0–25–50%), while careful control of tensile stress is needed to reduce the variability of resonance frequency for rigorous quantification. The effect of the tensile stress retained by the suspended DNA nanoresonators on the application of the technique is extensively explored using a combination of laser Doppler vibrometry and atomic force spectroscopy. DNA nanoresonators are real-time, label-free sensors and could avoid chemical functionalization and sample amplification. Therefore, they may represent in the future a key complementary routine tool for global DNA methylation analysis needed to evaluate the consequences of environmental stresses on the human genome.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Polyomavirus JCV excretion and genotype analysis in HIV-infected patients receiving highly active antiretroviral therapy

OBJECTIVE: To assess the frequency of shedding of polyomavirus JC virus (JCV) genotypes in urine of HIV-infected patients receiving highly active antiretroviral therapy (HAART). METHODS: Single samples of urine and blood were collected prospectively from 70 adult HIV-infected patients and 68 uninfected volunteers. Inclusion criteria for HIV-infected patients included an HIV RNA viral load < 1000 copies, CD4 cell count of 200-700 x 106 cells/l, and stable HAART regimen. PCR assays and sequence analysis were carried out using JCV-specific primers against different regions of the virus genome. RESULTS: JCV excretion in urine was more common in HIV-positive patients but not significantly different from that of the HIV-negative group [22/70 (31%) versus 13/68 (19%); P = 0.09]. HIV-positive patients lost the age-related pattern of JCV shedding (P = 0.13) displayed by uninfected subjects (P = 0.01). Among HIV-infected patients significant differences in JCV shedding were related to CD4 cell counts (P = 0.03). Sequence analysis of the JCV regulatory region from both HIV-infected patients and uninfected volunteers revealed all to be JCV archetypal strains. JCV genotypes 1 (36%) and 4 (36%) were the most common among HIV-infected patients, whereas type 2 (77%) was the most frequently detected among HIV-uninfected volunteers. CONCLUSION: These results suggest that JCV shedding is enhanced by modest depressions in immune function during HIV infection. JCV shedding occurred in younger HIV-positive persons than in the healthy controls. As the common types of JCV excreted varied among ethnic groups, JCV genotypes associated with progressive multifocal leukoencephalopathy may reflect demographics of those infected patient populations.

Non-NASA Center↗

The Analysis of the Patterns of Radiation-Induced DNA Damage Foci by a Stochastic Monte Carlo Model of DNA Double Strand Breaks Induction by Heavy Ions and Image Segmentation Software

To create a generalized mechanistic model of DNA damage in human cells that will generate analytical and image data corresponding to experimentally observed DNA damage foci and will help to improve the experimental foci yields by simulating spatial foci patterns and resolving problems with quantitative image analysis. Material and Methods: The analysis of patterns of RIFs (radiation-induced foci) produced by low- and high-LET (linear energy transfer) radiation was conducted by using a Monte Carlo model that combines the heavy ion track structure with characteristics of the human genome on the level of chromosomes. The foci patterns were also simulated in the maximum projection plane for flat nuclei. Some data analysis was done with the help of image segmentation software that identifies individual classes of RIFs and colocolized RIFs, which is of importance to some experimental assays that assign DNA damage a dual phosphorescent signal. Results: The model predicts the spatial and genomic distributions of DNA DSBs (double strand breaks) and associated RIFs in a human cell nucleus for a particular dose of either low- or high-LET radiation. We used the model to do analyses for different irradiation scenarios. In the beam-parallel-to-the-disk-of-a-flattened-nucleus scenario we found that the foci appeared to be merged due to their high density, while, in the perpendicular-beam scenario, the foci appeared as one bright spot per hit. The statistics and spatial distribution of regions of densely arranged foci, termed DNA foci chains, were predicted numerically using this model. Another analysis was done to evaluate the number of ion hits per nucleus, which were visible from streaks of closely located foci. In another analysis, our image segmentaiton software determined foci yields directly from images with single-class or colocolized foci. Conclusions: We showed that DSB clustering needs to be taken into account to determine the true DNA damage foci yield, which helps to determine the DSB yield. Using the model analysis, a researcher can refine the DSB yield per nucleus per particle. We showed that purely geometric artifacts, present in the experimental images, can be analytically resolved with the model, and that the quantization of track hits and DSB yields can be provided to the experimentalists who use enumeration of radiation-induced foci in immunofluorescence experiments using proteins that detect DNA damage. An automated image segmentaiton software can prove useful in a faster and more precise object counting for colocolized foci images.

Ponomarev, Artem↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Data for High Yield Production of 3-Hydroxypropionic Acid Using Issatchenkia orientalis

Biomanufacturing provides a more sustainable alternative to fossil-based chemical manufacturing. 3-Hydroxypropionic acid (3HP) is a top Department of Energy value-added chemical and precursor to bioplastics, yet cost-effective microbial production remains elusive. Here, we establish the acid-tolerant yeast Issatchenkia orientalis as a robust host for low-pH 3HP biosynthesis. Genome-scale modeling identifies the β-alanine pathway as optimal, offering the highest theoretical yield and lowest oxygen requirement. Thermodynamic analysis confirms its favorability under acidic conditions. Using sequence similarity network analysis, we discover highly active aspartate 1-decarboxylase (PAND), β-alanine-pyruvate aminotransferase (BAPAT), and 3HP dehydrogenase (YDFG), which significantly improve the pathway efficiency. Next, to further elevate the production, pathway optimization through multi-copy PAND integration, byproduct elimination (knockouts of pyruvate decarboxylase and glycerol-3-phosphate dehydrogenase), and reinforcement of aspartate flux by overexpression of pyruvate carboxylase and aspartate amino transferase improves the titer to 29 g/L in shake flasks. Fed-batch fermentation at pH 4 with low-cost corn steep liquor medium further increases the production to 92 g/L with 0.7 g/g yield and 0.55 g/L/h productivity. Techno-economic analysis indicates that such performance could potentially enable a financially viable process for sustainable acrylic acid production. This work establishes I. orientalis as a next-generation platform for cost-effective 3HP production and paves the way toward industrial commercialization.

Bioproducts↗

Developing a Genetic Variant Calling Pipeline for Quantifying the Complex Mutagenic Load Accumulated in BioNutrients-1 Production Pack Samples

Microorganisms hold great promise for on demand production of labile nutrients and pharmaceuticals as well recycling and in situ resource utilization. The utilization of microorganisms for such tasks on space missions is hindered by the limited data on how microbes respond to spaceflight. For example, the genetic stability of microorganisms, and the genomic engineered traits added to deliver desired functions, over long-term storage in the spacecraft environment is poorly understood. The BioNutrients-1 (BN-1) mission conducted a 5-year study of desiccated storage in Low Earth Orbit (LEO) to evaluate the suitability of eight synthetic biology chassis organisms for long-duration space missions. We are employing high-depth, whole genome sequencing (WGS) to determine the mutagenic load that accumulated during long-term storage. Mutation analysis pipelines are well established for homogenous culture grown from a single colony, but the mutational landscape of the BN-1 samples present a unique analysis challenge, as every cell in the BN-1 samples had a unique genetic journey of DNA damage and repair. Consequently, sequence variants are expected at low allele frequency within samples. To address this genetic complexity, we apply two distinct computational approaches to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO. For reference genome free mutation detection, we utilized DiscoSNP++, which is a de Bruijn graph approach. For reference genome-based mutation detection we utilize GATK for Microbes, which is a Bayesian probabilistic approach. We will benchmark these approaches against the mutations originally identified using CRISP, a method optimized for pooled samples. Ultimately, quantifying the mutation load imposed by storage or growth on the ISS will help identify chassis organisms with both high levels of genome stability and viability, which are desirable traits for implementation of bioproduction in long-duration missions.

SNP↗