Search NASA⌕ Search

SEARCH · Search NASA

Results for “gene presence/absence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

QTG-Finder2: A Generalized Machine-Learning Algorithm for Prioritizing QTL Causal Genes in Plants

Linkage mapping has been widely used to identify quantitative trait loci (QTL) in many plants and usually requires a time-consuming and labor-intensive fine mapping process to find the causal gene underlying the QTL. Previously, we described QTG-Finder, a machine-learning algorithm to rationally prioritize candidate causal genes in QTLs. Although it showed good performance, QTG-Finder could only be used in Arabidopsis and rice because of the limited number of known causal genes in other species. Here we tested the feasibility of enabling QTG-Finder to work on species that have few or no known causal genes by using orthologs of known causal genes as training set. The model trained with orthologs could recall about 64% of Arabidopsis and 83% of rice causal genes when the top 20% ranked genes were considered, which is similar to the performance of models trained with known causal genes. The average precision was 0.027 for Arabidopsis and 0.029 for rice. We further extended the algorithm to include polymorphisms in conserved non-coding sequences and gene presence/absence variation as additional features. Using this algorithm, QTG-Finder2, we trained and cross-validated Sorghum bicolor and Setaria viridis models. The S. bicolor model was validated by causal genes curated from the literature and could recall 70% of causal genes when the top 20% ranked genes were considered. Furthermore, we applied the S. viridis model and public transcriptome data to prioritize a plant height QTL and identified 13 candidate genes. QTL-Finder2 can accelerate the discovery of causal genes in any plant species and facilitate agricultural trait improvement.

59 BASIC BIOLOGICAL SCIENCES↗

Gradual polyploid genome evolution revealed by pan-genomic analysis of Brachypodium hybridum and its diploid progenitors

Abstract Our understanding of polyploid genome evolution is constrained because we cannot know the exact founders of a particular polyploid. To differentiate between founder effects and post polyploidization evolution, we use a pan-genomic approach to study the allotetraploid Brachypodium hybridum and its diploid progenitors. Comparative analysis suggests that most B. hybridum whole gene presence/absence variation is part of the standing variation in its diploid progenitors. Analysis of nuclear single nucleotide variants, plastomes and k-mers associated with retrotransposons reveals two independent origins for B. hybridum , ~1.4 and ~0.14 million years ago. Examination of gene expression in the younger B. hybridum lineage reveals no bias in overall subgenome expression. Our results are consistent with a gradual accumulation of genomic changes after polyploidization and a lack of subgenome expression dominance. Significantly, if we did not use a pan-genomic approach, we would grossly overestimate the number of genomic changes attributable to post polyploidization evolution.

59 BASIC BIOLOGICAL SCIENCES↗

De Novo Assembly and Annotation of 11 Diverse Shrub Willow (Salix) Genomes Reveals Novel Gene Organization in Sex-Linked Regions

Poplar and willow species in the Salicaceae are dioecious, yet have been shown to use different sex determination systems located on different chromosomes. Willows in the subgenus Vetrix are interesting for comparative studies of sex determination systems, yet genomic resources for these species are still quite limited. Only a few annotated reference genome assemblies are available, despite many species in use in breeding programs. Here we present de novo assemblies and annotations of 11 shrub willow genomes from six species. Copy number variation of candidate sex determination genes within each genome was characterized and revealed remarkable differences in putative master regulator gene duplication and deletion. We also analyzed copy number and expression of candidate genes involved in floral secondary metabolism, and identified substantial variation across genotypes, which can be used for parental selection in breeding programs. Lastly, we report on a genotype that produces only female descendants and identified gene presence/absence variation in the mitochondrial genome that may be responsible for this unusual inheritance.

59 BASIC BIOLOGICAL SCIENCES↗

BAD2matrix: Phylogenomic matrix concatenation, indel coding, and more

Common steps in phylogenomic matrix production include biological sequence concatenation, morphological data concatenation, insertion/deletion (indel) coding, gene content (presence/absence) coding, removing uninformative characters for parsimony analysis, recording with reduced amino acid alphabets, and occupancy filtering. Existing software does not accomplish these tasks on a phylogenomic scale using a single program. BAD2matrix is a Python script that performs the above-mentioned steps in phylogenomic matrix construction for DNA or amino acid sequences as well as morphological data. The script works in UNIX-like environments (e.g., LINUX, MacOS, Windows Subsystem for LINUX).

59 BASIC BIOLOGICAL SCIENCES↗

Adaptive gene loss in the common bean pan-genome during range expansion and domestication

The common bean ( Phaseolus vulgaris L.) is a crucial legume crop and an ideal evolutionary model to study adaptive diversity in wild and domesticated populations. Here, we present a common bean pan-genome based on five high-quality genomes and whole-genome reads representing 339 genotypes. It reveals ~234 Mb of additional sequences containing 6,905 protein-coding genes missing from the reference, constituting 49% of all presence/absence variants (PAVs). More non-synonymous mutations are found in PAVs than core genes, probably reflecting the lower effective population size of PAVs and fitness advantages due to the purging effect of gene loss. Our results suggest pan-genome shrinkage occurred during wild range expansion. Selection signatures provide evidence that partial or complete gene loss was a key adaptive genetic change in common bean populations with major implications for plant adaptation. The pan-genome is a valuable resource for food legume research and breeding for climate change mitigation and sustainable agriculture.

59 BASIC BIOLOGICAL SCIENCES↗

The full-length structure of Thermus scotoductus OLD defines the ATP hydrolysis properties and catalytic mechanism of Class 1 OLD family nucleases

Abstract OLD family nucleases contain an N-terminal ATPase domain and a C-terminal Toprim domain. Homologs segregate into two classes based on primary sequence length and the presence/absence of a unique UvrD/PcrA/Rep-like helicase gene immediately downstream in the genome. Although we previously defined the catalytic machinery controlling Class 2 nuclease cleavage, degenerate conservation of the C-termini between classes precludes pinpointing the analogous residues in Class 1 enzymes by sequence alignment alone. Our Class 2 structures also provide no information on ATPase domain architecture and ATP hydrolysis. Here we present the full-length structure of the Class 1 OLD nuclease from Thermus scotoductus (Ts) at 2.20 Å resolution, which reveals a dimerization domain inserted into an N-terminal ABC ATPase fold and a C-terminal Toprim domain. Structural homology with genome maintenance proteins identifies conserved residues responsible for Ts OLD ATPase activity. Ts OLD lacks the C-terminal helical domain present in Class 2 OLD homologs yet preserves the spatial organization of the nuclease active site, arguing that OLD proteins use a conserved catalytic mechanism for DNA cleavage. We also demonstrate that mutants perturbing ATP hydrolysis or DNA cleavage in vitro impair P2 OLD-mediated killing of recBC−Escherichia coli hosts, indicating that both the ATPase and nuclease activities are required for OLD function in vivo.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning enables identification of an alternative yeast galactose utilization pathway

How genomic differences contribute to phenotypic differences is a major question in biology. The recently characterized genomes, isolation environments, and qualitative patterns of growth on 122 sources and conditions of 1,154 strains from 1,049 fungal species (nearly all known) in the yeast subphylum Saccharomycotina provide a powerful, yet complex, dataset for addressing this question. We used a random forest algorithm trained on these genomic, metabolic, and environmental data to predict growth on several carbon sources with high accuracy. Known structural genes involved in assimilation of these sources and presence/absence patterns of growth in other sources were important features contributing to prediction accuracy. By further examining growth on galactose, we found that it can be predicted with high accuracy from either genomic (92.2%) or growth data (82.6%) but not from isolation environment data (65.6%). Prediction accuracy was even higher (93.3%) when we combined genomic and growth data. After the GALactose utilization genes, the most important feature for predicting growth on galactose was growth on galactitol, raising the hypothesis that several species in two orders, Serinales and Pichiales (containing the emerging pathogen Candida auris and the genus Ogataea, respectively), have an alternative galactose utilization pathway because they lack the GAL genes. Growth and biochemical assays confirmed that several of these species utilize galactose through an alternative oxidoreductive D-galactose pathway, rather than the canonical GAL pathway. Machine learning approaches are powerful for investigating the evolution of the yeast genotype–phenotype map, and their application will uncover novel biology, even in well-studied traits.

59 BASIC BIOLOGICAL SCIENCES↗

Single‐parent expression drives dynamic gene expression complementation in maize hybrids

SUMMARY Single‐parent expression (SPE) is defined as gene expression in only one of the two parents. SPE can arise from differential expression between parental alleles, termed non‐presence/absence (non‐PAV) SPE, or from the physical absence of a gene in one parent, termed PAV SPE. We used transcriptome data of diverse Zea mays (maize) inbreds and hybrids, including 401 samples from five different tissues, to test for differences between these types of SPE genes. Although commonly observed, SPE is highly genotype and tissue specific. A positive correlation was observed between the genetic distance of the two inbred parents and the number of SPE genes identified. Regulatory analysis showed that PAV SPE and non‐PAV SPE genes are mainly regulated by cis effects, with a small fraction under trans regulation. Polymorphic transposable element insertions in promoter sequences contributed to the high level of cis regulation for PAV SPE and non‐PAV SPE genes. PAV SPE genes were more frequently expressed in hybrids than non‐PAV SPE genes. The expression of parentally silent alleles in hybrids of non‐PAV SPE genes was relatively rare but occurred in most hybrids. Non‐PAV SPE genes with expression of the silent allele in hybrids are more likely to exhibit above high parent expression level than hybrids that do not express the silent allele, leading to non‐additive expression. This study provides a comprehensive understanding of the nature of non‐PAV SPE and PAV SPE genes and their roles in gene expression complementation in maize hybrids.

Li, Zhi↗

Metabolic Source Isotopic Pair Labeling and Genome-Wide Association Are Complementary Tools for the Identification of Metabolite-Gene Associations in Plants

The optimal extraction of information from untargeted metabolomics analyses is a continuing challenge. Here, we describe an approach that combines stable isotope labeling, liquid chromatography– mass spectrometry (LC–MS), and a computational pipeline to automatically identify metabolites produced from a selected metabolic precursor. Here, we identified the subset of the soluble metabolome generated from phenylalanine (Phe) in Arabidopsis thaliana, which we refer to as the Phe-derived metabolome (FDM) In addition to identifying Phe-derived metabolites present in a single wild-type reference accession, the FDM was established in nine enzymatic and regulatory mutants in the phenylpropanoid pathway. To identify genes associated with variation in Phe-derived metabolites in Arabidopsis, MS features collected by untargeted metabolite profiling of an Arabidopsis diversity panel were retrospectively annotated to the FDM and natural genetic variants responsible for differences in accumulation of FDM features were identified by genome-wide association. Large differences in Phe-derived metabolite accumulation and presence/absence variation of abundant metabolites were observed in the nine mutants as well as between accessions from the diversity panel. Many Phe-derived metabolites that accumulated in mutants also accumulated in non-Col-0 accessions and was associated to genes with known or suspected functions in the phenylpropanoid pathway as well as genes with no known functions. Overall, we show that cataloguing a biochemical pathway’s products through isotopic labeling across genetic variants can substantially contribute to the identification of metabolites and genes associated with their biosynthesis.

09 BIOMASS FUELS↗

Genetics and Genomics of Pathogen Resistance in Switchgrass (Final Report)

This project was funded by DOE under Grant no. DE-SC0016108. Originally approved for the 2016-2019 period, two no-cost extensions were solicited and approved, which prolonged the lifespan through July 2021. This final report informs on the results obtained so far from the research implemented. The research hinged on integrating genomics (genomic selection, RNAseq, virus-plant interactions) with classical genetics (conventional breeding) to incorporate durable resistance to fungal (rust) and viral (mosaic) diseases in switchgrass (Panicum virgatum) populations being bred for bioenergy. Higher biomass yield, higher quality (low lignin content), and durable disease resistance are key features to make lignocellulosic switchgrass feedstocks economically competitive and sustainable. Genomic selection is being applied on three generations of a switchgrass population derived from crossing two ecotypes (Kanlow as lowland female and Summer as upland male) with differential performance in terms of biomass yield and quality, disease resistance, and winter survivability. Target populations were screened for rust and mosaic in field and/or lab and phenotyped for biomass yield and quality traits. Genetic analyses were applied across generations to capture the joint inheritance of the targeted traits and predict breeding values for parents and progeny with greater accuracy. Parental and a panel of different switchgrass populations were genotyped with the DArTseq technology to develop SNP (0, 1, 2) and in-silico (presence/absence) DArT markers. Rust inoculations techniques were developed and applied successfully on switchgrass. The original populations (Kanlow and Summer) were sequenced with RNAseq to capture the gene expression profiles across sequential time-points and appraise the basis of greater resistance in the Kanlow vs the Summer ecotype. Constructs of PMV and sPMV mosaic virus were assembled and tested first on proso millet to find the best protocol to use later on switchgrass. Results from the preliminary analyses indicate that 1) ample additive genetic variation is available for selection and improving this inter-ecotypic population for yield, quality, and disease traits, 2) significant gains are to be expected with the genetic correlations being favorable between yield and lignin content and between yield and disease ratings, 3) substantial differences exist in the genetic regions controlling rust resistance in the two ecotypes, 4) co-infection with PMV isolates from Nebraska and its satellite from Kansas elicit severe mosaic symptoms, and 5) two different genetic systems are responsible for imparting resistance to rust and virus in switchgrass.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of Harvest on Switchgrass Leaf Microbial Communities

Switchgrass is a promising feedstock for biofuel production, with potential for leveraging its native microbial community to increase productivity and resilience to environmental stress. Here, we characterized the bacterial, archaeal and fungal diversity of the leaf microbial community associated with four switchgrass (Panicum virgatum) genotypes, subjected to two harvest treatments (annual harvest and unharvested control), and two fertilization levels (fertilized and unfertilized control), based on 16S rRNA gene and internal transcribed spacer (ITS) region amplicon sequencing. Leaf surface and leaf endosphere bacterial communities were significantly different with Alphaproteobacteria enriched in the leaf surface and Gammaproteobacteria and Bacilli enriched in the leaf endosphere. Harvest treatment significantly shifted presence/absence and abundances of bacterial and fungal leaf surface community members: Gammaproteobacteria were significantly enriched in harvested and Alphaproteobacteria were significantly enriched in unharvested leaf surface communities. These shifts were most prominent in the upland genotype DAC where the leaf surface showed the highest enrichment of Gammaproteobacteria, including taxa with 100% identity to those previously shown to have phytopathogenic function. Fertilization did not have any significant impact on bacterial or fungal communities. We also identified bacterial and fungal taxa present in both the leaf surface and leaf endosphere across all genotypes and treatments. These core taxa were dominated by Methylobacterium, Enterobacteriaceae, and Curtobacterium, in addition to Aureobasidium, Cladosporium, Alternaria and Dothideales. Local core leaf bacterial and fungal taxa represent promising targets for plant microbe engineering and manipulation across various genotypes and harvest treatments. Our study showcases, for the first time, the significant impact that harvest treatment can have on bacterial and fungal taxa inhabiting switchgrass leaves and the need to include this factor in future plant microbial community studies.

59 BASIC BIOLOGICAL SCIENCES↗

Virulence factors and antimicrobial resistance profiles of Campylobacter isolates recovered from consecutively reused broiler litter

ABSTRACT Campylobacterinfections are a leading cause of bacterial diarrhea in humans globally. Infections are due to consumption of contaminated food products and are highly associated with chicken meat, with chickens being an important reservoir forCampylobacter. Here, we characterized the genetic diversity ofCampylobacter jejuni(C. jejuni) andCampylobacter coli(C. coli) detected in broiler chicken litter over three consecutive flocks and determined their antimicrobial resistance (ARM) and virulence factor (VF) profiles.Campylobacterwas detected in 9.38% (27/288) of litter samples collected. Antimicrobial susceptibility testing and whole genome sequencing were performed onC. jejuni(n= 39) andC. coli(n= 5) isolates.Campylobactervirulence factors differed within and across broiler houses but were explained by the broiler flock cohort raised on litter,Campylobacterspecies andCampylobactermultilocus sequence type (MLST). Virulence factors involved in the ability to invade and colonize host tissues and evade host defenses were present inC. jejuniisolates (ST-464) from flock cohorts 1 and 2 but absent inC. jejuniisolates (ST-48) from flock cohort 3.C. jejuniisolates from house three harbored a significantly higher proportion of virulence genes with functions related to glycosylation and immune evasion thanC. jejuniisolates from houses 1 and 2 (P< 0.01). AllC. jejuniisolates were susceptible to all antibiotics tested whileC. coli(n= 4) were resistant to tetracycline and harbored the tetracycline resistant ribosomal protection protein (TetO). Our results suggest that house environment and broiler management practices imposed selective pressures on virulence factors and antimicrobial resistance genes ofCampylobacter. IMPORTANCE Campylobacteris a leading cause of foodborne illness in the United States due to consumption of contaminated or mishandled food products, often associated with chicken meat.Campylobacteris common in the microbiota of avian and mammalian gut; however, acquisition of antimicrobial resistance genes (ARGs) and virulence factors (VFs) may result in strains that pose significant threat to public health. Although there are studies investigating the genetic diversity ofCampylobacterstrains isolated from post-harvest chicken samples, there are limited data on the genome characteristics of isolates recovered from preharvest broiler production. Here, we show thatCampylobacter jejuniandCampylobacter colidiffer in their carriage of antimicrobial resistance and virulence factors may also differ in their ability to persist in litter during consecutive grow-out of broiler flocks. We found that presence/absence of virulence factors needed for evasion of host defense mechanisms and gut colonization played an integral role in differentiatingCampylobacterstrains.

Microbiology↗

Two major chromosome evolution events with unrivaled conserved gene content in pomegranate

Pomegranate has a unique evolutionary history given that different cultivars have eight or nine bivalent chromosomes with possible crossability between the two classes. Therefore, it is important to study chromosome evolution in pomegranate to understand the dynamics of its population. Here, we de novo assembled the Azerbaijani cultivar “Azerbaijan guloyshasi” (AG2017; 2n = 16) and re-sequenced six cultivars to track the evolution of pomegranate and to compare it with previously published de novo assembled and re-sequenced cultivars. High synteny was observed between AG2017, Bhagawa (2n = 16), Tunisia (2n = 16), and Dabenzi (2n = 18), but these four cultivars diverged from the cultivar Taishanhong (2n = 18) with several rearrangements indicating the presence of two major chromosome evolution events. Major presence/absence variations were not observed as >99% of the five genomes aligned across the cultivars, while >99% of the pan-genic content was represented by Tunisia and Taishanhong only. We also revisited the divergence between soft- and hard-seeded cultivars with less structured population genomic data, compared to previous studies, to refine the selected genomic regions and detect global migration routes for pomegranate. We reported a unique admixture between soft- and hard-seeded cultivars that can be exploited to improve the diversity, quality, and adaptability of local pomegranate varieties around the world. Our study adds body knowledge to understanding the evolution of the pomegranate genome and its implications for the population structure of global pomegranate diversity, as well as planning breeding programs aiming to develop improved cultivars.

59 BASIC BIOLOGICAL SCIENCES↗

Xylella fastidiosa causes transcriptional shifts that precede tylose formation and starch depletion in xylem

Abstract Pierce's disease (PD) in grapevine ( Vitis vinifera ) is caused by the bacterial pathogen Xylella fastidiosa . X. fastidiosa is limited to the xylem tissue and following infection induces extensive plant‐derived xylem blockages, primarily in the form of tyloses. Tylose‐mediated vessel occlusions are a hallmark of PD, particularly in susceptible V. vinifera . We temporally monitored tylose development over the course of the disease to link symptom severity to the level of tylose occlusion and the presence/absence of the bacterial pathogen at fine‐scale resolution. The majority of vessels containing tyloses were devoid of bacterial cells, indicating that direct, localized perception of X. fastidiosa was not a primary cause of tylose formation. In addition, we used X‐ray computed microtomography and machine‐learning to determine that X. fastidiosa induces significant starch depletion in xylem ray parenchyma cells. This suggests that a signalling mechanism emanating from the vessels colonized by bacteria enables a systemic response to X. fastidiosa infection. To understand the transcriptional changes underlying these phenotypes, we integrated global transcriptomics into the phenotypes we tracked over the disease spectrum. Differential gene expression analysis revealed that considerable transcriptomic reprogramming occurred during early PD before symptom appearance. Specifically, we determined that many genes associated with tylose formation (ethylene signalling and cell wall biogenesis) and drought stress were up‐regulated during both Phase I and Phase II of PD. On the contrary, several genes related to photosynthesis and carbon fixation were down‐regulated during both phases. These responses correlate with significant starch depletion observed in ray cells and tylose synthesis in vessels.

59 BASIC BIOLOGICAL SCIENCES↗