Search NASASearch

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Populus PtrbHLH011 Is a Transcriptional Co‐Regulator Involved in the Activation of Cell Wall Biosynthesis by Iron Deprivation

The lack of a mechanistic understanding of the environmental plasticity of secondary cell wall (SCW) biosynthesis restricts large-scale biomass and bioenergy production on marginal lands. Using Populus (poplar), a key bioenergy crop, we discovered that iron deprivation, a prevalent abiotic stress on marginal lands, stimulates SCW biosynthesis in stems. We identified the transcription factor PtrbHLH011 as a critical regulator underlying this response. Through integrated analyses involving phenotypic characterisation of PtrbHLH011 knockout and overexpression plants, functional genomics and molecular investigations, we established that PtrbHLH011 functions as a central regulator of SCW biosynthesis, iron homeostasis and flavonoid biosynthesis by directly repressing essential genes in these pathways. Iron deprivation downregulates PtrbHLH011 expression, subsequently activating these biosynthetic pathways. Notably, cytosine base editing-based knockout of PtrbHLH011 significantly enhanced plant growth, yielding up to a 110% increase in stem diameter and a 300% increase in leaf iron content. These findings present a novel regulatory mechanism linking environmental iron availability to SCW biosynthesis and illustrate a practical strategy to improve biomass yield on iron-deficient marginal lands. Furthermore, our mechanistic insights into PtrbHLH011 target recognition and regulation provide a valuable foundation for precise manipulation of gene regulatory networks, facilitating the development of high-performance bioenergy crops adapted to marginal environments.

59 BASIC BIOLOGICAL SCIENCES

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)

High phenotypic and genotypic plasticity among strains of the mushroom-forming fungus Schizophyllum commune

Schizophyllum commune is a mushroom-forming fungus notable for its distinctive fruiting bodies with split gills. It is used as a model organism to study mushroom development, lignocellulose degradation and mating type loci. It is a hypervariable species with considerable genetic and phenotypic diversity between the strains. In this study, we systematically phenotyped 16 dikaryotic strains for aspects of mushroom development and 18 monokaryotic strains for lignocellulose degradation. There was considerable heterogeneity among the strains regarding these phenotypes. The majority of the strains developed mushrooms with varying morphologies, although some strains only grew vegetatively under the tested conditions. Growth on various carbon sources showed strain-specific profiles. The genomes of seven monokaryotic strains were sequenced and analyzed together with six previously published genome sequences. Moreover, the related species Schizophyllum fasciatum was sequenced. Although there was considerable genetic variation between the genome assemblies, the genes related to mushroom formation and lignocellulose degradation were well conserved. These sequenced genomes, in combination with the high phenotypic diversity, will provide a solid basis for functional genomics analyses of the strains of S. commune.

59 BASIC BIOLOGICAL SCIENCES

Through the lens of bioenergy crops: advances, bottlenecks, and promises of plant engineering

Advances in engineering of bioenergy crops were driven over the past years by adapting technological breakthroughs and accelerating conventional applications but also exposed intriguing challenges. New tools revealed rich interconnectivity in the exponentially growing and dynamic 'big' omics data' of metabolomes, transcriptomes, and genomes at previously inaccessible magnitude (global, cross-species, meta-) and resolution (single cell). Insights enabled fresh hypotheses and stimulated disciplines such as functional genomics with discovery of broad regulatory networks and their determinants, that is, DNA parts, including promoters, regulatory elements, and transcription factors. Their rational design, assembly into increasingly complex blueprints, and installation into diverse chassis is an existing frontier that may benefit from emerging technologies to address bottlenecks. Interweaving nature-inspired to fully synthetic parts has already allowed building of fine-tuned regulatory circuits, or new-to-nature metabolic routes insulated from the biological context of the chassis species. Similarly, developments and the evolving need for unifying principles in plant transformation and species-agnostic technologies highlight future opportunities for engineering the next generation of bioenergy plants.

60 APPLIED LIFE SCIENCES

Supporting Information for manuscript: “A latitudinal gradient in S/G lignin monomer ratio driven by laccase in natural poplar variants”

Lignin composition plays a crucial role in plant structural integrity and environmental adaptation. However, the genetic and molecular mechanisms underlying natural variation in lignin composition remain poorly understood. This study investigates the syringyl-to-guaiacyl (S/G) lignin monomer ratio across a natural population of Populus trichocarpa spanning a latitudinal gradient along the Northwest coast of North America. By integrating biochemical, genomic, and geographic analysis, we identify key gene variants associated with S/G ratio differences. These datasets provide valuable insights into the evolutionary and functional genomics of lignin composition and serve as a resource for developing poplar variants optimized for forestry and bioenergy applications.

Poplar, lignin composition, laccases, latitude, ad

BSMV-mediated genome editing exhibits host-specific heritability: germline transmission in barley and somatic edits in Nicotiana benthamiana

Plant RNA virus–mediated guide RNA (gRNA) delivery represents a transformative advance in genome editing technologies. Unlike conventional transformation methods that rely on labor-intensive tissue culture and regeneration for each individual gRNA delivery, viral vectors can rapidly and systemically transmit gRNAs into pre-established Cas-expressing plants, providing an accelerated route for functional genomics and trait discovery directly in planta . However, key design parameters, including subgenomic promoter choice, transcript architecture, and their effects on viral fitness and editing outcomes, remain to be elucidated for most viral platforms. We developed five Barley stripe mosaic virus (BSMV) vectors, each with distinct subgenomic promoter elements to drive single gRNA expression. These were initially evaluated in Cas9-expressing transgenic Nicotiana benthamiana plants targeting the Phytoene desaturase ( PDS ) gene to compare their editing efficiencies. Single gRNAs expressed under the duplicated γb subgenomic promoter or when fused directly to the γb genome achieved the highest mutation frequencies (up to 90% at 60 days post-inoculation), whereas β1- and β2-driven sgRNAs produced delayed and reduced editing. Thus, promoter selection critically determines gRNA accumulation and the efficacy of BSMV-mediated genome editing. The top-performing design was then applied to Cas9-expressing barley ( Hordeum vulgare ) targeting HvCMF7 (conferring green-white variegation) and HvGW2.1 (impacts grain width and weight). BSMV spread systemically throughout barley, inducing somatic and heritable mutations at frequencies up to 100%, with virus-free edited progeny. In contrast, despite robust somatic editing in N. benthamiana, no heritable mutations were detected indicating species-dependent limitations in germline transmission. Our systematic comparison of subgenomic promoter architectures establishes clear design principles for optimizing viral vector–mediated delivery. Promoter choice and transcript structure critically shape editing efficiency and viral stability. The host-specific boundary for germline editing, defined by efficient heritable editing in barley but not N. benthamiana , highlights where BSMV offers advantages and where alternative vectors or hybrid strategies are required, guiding rational platform selection for diverse crop species and applications. Collectively, these findings establish BSMV as a promising next-generation vector for rapid, tissue culture–free, and transformation-independent genome editing in cereals and other recalcitrant monocots.

barley

Soybean genomics research community strategic plan: A vision for 2024–2028

Abstract This strategic plan summarizes the major accomplishments achieved in the last quinquennial by the soybean [Glycine max(L.) Merr.] genetics and genomics research community and outlines key priorities for the next 5 years (2024–2028). This work is the result of deliberations among over 50 soybean researchers during a 2‐day workshop in St Louis, MO, USA, at the end of 2022. The plan is divided into seven traditional areas/disciplines: Breeding, Biotic Interactions, Physiology and Abiotic Stress, Functional Genomics, Biotechnology, Genomic Resources and Datasets, and Computational Resources. One additional section was added, Training the Next Generation of Soybean Researchers, when it was identified as a pressing issue during the workshop. This installment of the soybean genomics strategic plan provides a snapshot of recent progress while looking at future goals that will improve resources and enable innovation among the community of basic and applied soybean researchers. We hope that this work will inform our community and increase support for soybean research.

Genetics & Heredity

Diverse signatures of convergent evolution in cactus-associated yeasts

Many distantly related organisms have convergently evolved traits and lifestyles that enable them to live in similar ecological environments. However, the extent of phenotypic convergence evolving through the same or distinct genetic trajectories remains an open question. Here, we leverage a comprehensive dataset of genomic and phenotypic data from 1,049 yeast species in the subphylum Saccharomycotina (Kingdom Fungi, Phylum Ascomycota) to explore signatures of convergent evolution in cactophilic yeasts, ecological specialists associated with cacti. We inferred that the ecological association of yeasts with cacti arose independently approximately 17 times. Using a machine learning–based approach, we further found that cactophily can be predicted with 76% accuracy from both functional genomic and phenotypic data. The most informative feature for predicting cactophily was thermotolerance, which we found to be likely associated with altered evolutionary rates of genes impacting the cell envelope in several cactophilic lineages. We also identified horizontal gene transfer and duplication events of plant cell wall–degrading enzymes in distantly related cactophilic clades, suggesting that putatively adaptive traits evolved independently through disparate molecular mechanisms. Notably, we found that multiple cactophilic species and their close relatives have been reported as emerging human opportunistic pathogens, suggesting that the cactophilic lifestyle—and perhaps more generally lifestyles favoring thermotolerance—might preadapt yeasts to cause human disease. This work underscores the potential of a multifaceted approach involving high-throughput genomic and phenotypic data to shed light onto ecological adaptation and highlights how convergent evolution to wild environments could facilitate the transition to human pathogenicity.

59 BASIC BIOLOGICAL SCIENCES

RNAi and genome editing of sugarcane: Progress and prospects

SUMMARY Sugarcane, which provides 80% of global table sugar and 40% of biofuel, presents unique breeding challenges due to its highly polyploid, heterozygous, and frequently aneuploid genome. Significant progress has been made in developing genetic resources, including the recently completed reference genome of the sugarcane cultivar R570 and pan‐genomic resources from sorghum, a closely related diploid species. Biotechnological approaches including RNA interference (RNAi), overexpression of transgenes, and gene editing technologies offer promising avenues for accelerating sugarcane improvement. These methods have successfully targeted genes involved in important traits such as sucrose accumulation, lignin biosynthesis, biomass oil accumulation, and stress response. One of the main transformation methods—biolistic gene transfer or Agrobacterium ‐mediated transformation—coupled with efficient tissue culture protocols, is typically used for implementing these biotechnology approaches. Emerging technologies show promise for overcoming current limitations. The use of morphogenic genes can help address genotype constraints and improve transformation efficiency. Tissue culture‐free technologies, such as spray‐induced gene silencing, virus‐induced gene silencing, or virus‐induced gene editing, offer potential for accelerating functional genomics studies. Additionally, novel approaches including base and prime editing, orthogonal synthetic transcription factors, and synthetic directed evolution present opportunities for enhancing sugarcane traits. These advances collectively aim to improve sugarcane's efficiency as a crop for both sugar and biofuel production. This review aims to discuss the progress made in sugarcane methodologies, with a focus on RNAi and gene editing approaches, how RNAi can be used to inform functional gene targets, and future improvements and applications.

Brant, Eleanor [Agronomy Department, Plant Molecul

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing

Targeted seed EMS mutagenesis reveals a basic helix–loop–helix transcription factor underlying male sterility in sorghum

Abstract Forward genetic screens of mutant populations are fundamental for functional genomics studies. However, isolating independent mutant alleles to molecularly identify causal genes is challenging in species recalcitrant to genetic manipulation. Here, we demonstrate that classic seed ethyl methanesulfonate (EMS) mutagenesis coupled with genome sequencing can overcome this limitation in sorghum. We used this method to generate new mutant alleles of sorghum MALE STERILE 8 (MS8) and identified the causal locus for the ms8 phenotype as Sobic.004G270900, which encodes the sorghum ortholog of maize bhlh122, a basic helix–loop–helix (bHLH) transcription factor required for male fertility in maize. Bulked segregant analysis mapped ms8-1 to a region on chromosome 4 containing Sobic.004G270900. Seeds from heterozygous MS8/ms8-1 plants were mutagenized and screened for chimeric inflorescences containing sectors with white, sterile anthers resembling the ms8-1 homozygous phenotype. DNA sequencing of sterile and fertile sectors from a single chimeric inflorescence revealed two mutations in Sobic.004G270900 within the sterile sector, but not the fertile sector. Isolation of this loss-of-function allele (ms8-2) established Sobic.004G270900 as the causative locus for male sterility in the ms8 mutant. We generated additional alleles of MS8 in a different genetic background using CRISPR/Cas9-based gene editing, where deletions in Sobic.004G270900 also resulted in male sterility. Our work identified a gene underlying male sterility in sorghum and provides a novel and straightforward genetic tool for researchers who lack access to advanced transformation facilities to validate gene candidates. Unlike gene editing, no prior knowledge of candidate genes is required for targeted seed EMS mutagenesis to aid identification of causal loci.

Genetics & Heredity

Rapid and efficient in planta genome editing in sorghum using foxtail mosaic virus‐mediated sgRNA delivery

SUMMARY The requirement of in vitro tissue culture for the delivery of gene editing reagents limits the application of gene editing to commercially relevant varieties of many crop species. To overcome this bottleneck, plant RNA viruses have been deployed as versatile tools for in planta delivery of recombinant RNA. Viral delivery of single‐guide RNAs (sgRNAs) to transgenic plants that stably express CRISPR‐associated (Cas) endonuclease has been successfully used for targeted mutagenesis in several dicotyledonous and few monocotyledonous plants. Progress with this approach in monocotyledonous plants is limited so far by the availability of effective viral vectors. We engineered a set of foxtail mosaic virus (FoMV) and barley stripe mosaic virus (BSMV) vectors to deliver the fluorescent protein AmCyan to track viral infection and movement in Sorghum bicolor . We further used these viruses to deliver and express sgRNAs to Cas9 and Green Fluorescent Protein (GFP) expressing transgenic sorghum lines, targeting Phytoene desaturase ( PDS ), Magnesium‐chelatase subunit I ( MgCh ), 4‐hydroxy‐3‐methylbut‐2‐enyl diphosphate reductase , orthologs of maize Lemon white1 ( Lw1 ) or GFP . The recombinant BSMV did neither infect sorghum nor deliver or express AmCyan and sgRNAs. In contrast, the recombinant FoMV systemically spread throughout sorghum plants and induced somatic mutations with frequencies reaching up to 60%. This mutagenesis led to visible phenotypic changes, demonstrating the potential of FoMV for in planta gene editing and functional genomics studies in sorghum.

54 ENVIRONMENTAL SCIENCES

Data for Rapid and Efficient in planta Genome Editing in Sorghum Using Foxtail Mosaic Virus-mediated sgRNA Delivery

The requirement of in vitro tissue culture for the delivery of gene editing reagents limits the application of gene editing to commercially relevant varieties of many crop species. To overcome this bottleneck, plant RNA viruses have been deployed as versatile tools for in planta delivery of recombinant RNA. Viral delivery of single-guide RNAs (sgRNAs) to transgenic plants that stably express CRISPR-associated (Cas) endonuclease has been successfully used for targeted mutagenesis in several dicotyledonous and few monocotyledonous plants. Progress with this approach in monocotyledonous plants is limited so far by the availability of effective viral vectors. We engineered a set of foxtail mosaic virus (FoMV) and barley stripe mosaic virus (BSMV) vectors to deliver the fluorescent protein AmCyan to track viral infection and movement in Sorghum bicolor . We further used these viruses to deliver and express sgRNAs to Cas9 and Green Fluorescent Protein (GFP) expressing transgenic sorghum lines, targeting Phytoene desaturase (PDS), Magnesium-chelatase subunit I (MgCh), 4-hydroxy-3-methylbut-2-enyl diphosphate reductase, orthologs of maize Lemon white1 (Lw1) or GFP. The recombinant BSMV did neither infect sorghum nor deliver or express AmCyan and sgRNAs. In contrast, the recombinant FoMV systemically spread throughout sorghum plants and induced somatic mutations with frequencies reaching up to 60%. This mutagenesis led to visible phenotypic changes, demonstrating the potential of FoMV for in planta gene editing and functional genomics studies in sorghum.

Feedstock Production

Learning genetic perturbation effects with variational causal inference

Advances in sequencing technologies have enhanced the understanding of gene regulation in cells. In particular, Perturb-seq has enabled high-resolution profiling of the transcriptomic response to genetic perturbations at the single-cell level. This understanding has implications in functional genomics and potentially for identifying therapeutic targets. Various computational models have been developed to predict perturbational effects. While deep learning models excel at interpolating observed perturbational data, they tend to overfit in the lack of enough data and may not generalize well to unseen perturbations. In contrast, mechanistic models, such as linear causal models based on gene regulatory networks, hold greater potential for extrapolation, as they encapsulate regulatory information that can predict responses to unseen perturbations. However, their application has been limited to small studies due to overly simplistic assumptions, making them less effective in handling noisy, large-scale single-cell data. We propose a hybrid approach that combines a mechanistic causal model with variational deep learning, termed Single Cell Causal Variational Autoencoder (SCCVAE). The mechanistic model employs a learned regulatory network to represent perturbational changes as shift interventions that propagate through the learned network. SCCVAE integrates this mechanistic causal model into a variational autoencoder, generating rich, comprehensive transcriptomic responses. Our results indicate that SCCVAE exhibits superior performance over current state-of-the-art baselines for extrapolating to predict unseen perturbational responses. Additionally, for the observed perturbations, the latent space learned by SCCVAE allows for the identification of functional perturbation modules and simulation of single-gene knockdown experiments of varying penetrance, presenting a robust tool for interpreting and interpolating perturbational responses at the single-cell level.

59 BASIC BIOLOGICAL SCIENCES

Unique trajectory of gene family evolution from genomic analysis of nearly all known species in an ancient yeast lineage

Gene gains and losses are a major driver of genome evolution; their precise characterization can provide insights into the origin and diversification of major lineages. Here, we examined gene family evolution of 1154 genomes from nearly all known species in the medically and technologically important yeast subphylum Saccharomycotina. We found that yeast gene family evolution differs from that of plants, animals, and filamentous ascomycetes, and is characterized by smaller overall gene numbers yet larger gene family sizes for a given gene number. Faster-evolving lineages (FELs) in yeasts experienced significantly higher rates of gene losses—commensurate with a narrowing of metabolic niche breadth—but higher speciation rates than their slower-evolving sister lineages (SELs). Gene families most often lost are those involved in mRNA splicing, carbohydrate metabolism, and cell division and are likely associated with intron loss, metabolic breadth, and non-canonical cell cycle processes. Our results highlight the significant role of gene family contractions in the evolution of yeast metabolism, genome function, and speciation, and suggest that gene family evolutionary trajectories have differed markedly across major eukaryotic lineages.

Comparative Genomics

Omics-driven onboarding of the carotenoid producing red yeast Xanthophyllomyces dendrorhous CBS 6938

Transcriptomics is a powerful approach for functional genomics and systems biology, yet it can also be used for genetic part discovery. Here, we derive constitutive and light-regulated promoters directly from transcriptomics data of the basidiomycete red yeast Xanthophyllomyces dendrorhous CBS 6938 (anamorph Phaffia rhodozyma) and use these promoters with other genetic elements to create a modular synthetic biology parts collection for this organism. X. dendrorhous is currently the sole biotechnologically relevant yeast in the Tremellomycete class-it produces large amounts of astaxanthin, especially under oxidative stress and exposure to light. Thus, we performed transcriptomics on X. dendrorhous under different wavelengths of light (red, green, blue, and ultraviolet) and oxidative stress. Differential gene expression analysis (DGE) revealed that terpenoid biosynthesis was primarily upregulated by light through crtI, while oxidative stress upregulated several genes in the pathway. Further gene ontology (GO) analysis revealed a complex survival response to ultraviolet (UV) where X. dendrorhous upregulates aromatic amino acid and tetraterpenoid biosynthesis and downregulates central carbon metabolism and respiration. The DGE data was also used to identify 26 constitutive and regulated genes, and then, putative promoters for each of the 26 genes were derived from the genome. Simultaneously, a modular cloning system for X. dendrorhous was developed, including integration sites, terminators, selection markers, and reporters. Each of the 26 putative promoters were integrated into the genome and characterized by luciferase assay in the dark and under UV light. The putative constitutive promoters were constitutive in the synthetic genetic context, but so were many of the putative regulated promoters. Notably, one putative promoter, derived from a hypothetical gene, showed ninefold activation upon UV exposure. Thus, this study reveals metabolic pathway regulation and develops a genetic parts collection for X. dendrorhous from transcriptomic data. Therefore, this study demonstrates that combining systems biology and synthetic biology into an omics-to-parts workflow can simultaneously provide useful biological insight and genetic tools for nonconventional microbes, particularly those without a related model organism. This approach can enhance current efforts to engineer diverse microbes.

60 APPLIED LIFE SCIENCES

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics