Search NASA⌕ Search

Engineering topics

Clark, Lindsay V.

Publications and source records attributed to Clark, Lindsay V..

Improving precision and accuracy of genetic mapping with genotyping-by-sequencing data in outcrossing species

This dataset contains all data and supplementary materials from "Improving precision and accuracy of genetic mapping with genotyping-by-sequencing data in outcrossing species". An Excel file a list of all QTLs and linkage group length (in cM) obtained with two different SNP-calling methods (Tassel-Uneak and Tassel-GBS), genetic map-construction method (linkage-only and reference order-corrected) and depth filters (12x, 20x, 30x and 40x) for genetic mapping of 18 biomass yield traits in a biparental Miscanthus sinensis population using RAD-Seq SNPs is provided as "Supplementary file 1". A Perl script with the code for filtering VCF and HapMap-formatted data files is provided as “Supplementary file 2”. Phenotype data used for QTL mapping is provided as “Supplementary File 3”. A Perl script with the code for the simulation study is provided as “Supplementary file 4”.

GenotypingSimulator↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of genotype‐calling methodologies on genome‐wide association and genomic prediction in polyploids

Abstract Discovery and analysis of genetic variants underlying agriculturally important traits are key to molecular breeding of crops. Reduced representation approaches have provided cost‐efficient genotyping using next‐generation sequencing. However, accurate genotype calling from next‐generation sequencing data is challenging, particularly in polyploid species due to their genome complexity. Recently developed Bayesian statistical methods implemented in available software packages, polyRAD, EBG, and updog, incorporate error rates and population parameters to accurately estimate allelic dosage across any ploidy. We used empirical and simulated data to evaluate the three Bayesian algorithms and demonstrated their impact on the power of genome‐wide association study (GWAS) analysis and the accuracy of genomic prediction. We further incorporated uncertainty in allelic dosage estimation by testing continuous genotype calls and comparing their performance to discrete genotypes in GWAS and genomic prediction. We tested the genotype‐calling methods using data from two autotetraploid species, Miscanthus sacchariflorus and Vaccinium corymbosum , and performed GWAS and genomic prediction. In the empirical study, the tested Bayesian genotype‐calling algorithms differed in their downstream effects on GWAS and genomic prediction, with some showing advantages over others. Through subsequent simulation studies, we observed that at low read depth, polyRAD was advantageous in its effect on GWAS power and limit of false positives. Additionally, we found that continuous genotypes increased the accuracy of genomic prediction, by reducing genotyping error, particularly at low sequencing depth. Our results indicate that by using the Bayesian algorithm implemented in polyRAD and continuous genotypes, we can accurately and cost‐efficiently implement GWAS and genomic prediction in polyploid crops.

59 BASIC BIOLOGICAL SCIENCES↗

Host genetic variation drives the differentiation in the ecological role of the native Miscanthus root-associated microbiome

Microbiome recruitment is influenced by plant host, but how host plant impacts the assembly, functions, and interactions of perennial plant root microbiomes is poorly understood. Here we examined prokaryotic and fungal communities between rhizosphere soils and the root endophytic compartment in two native Miscanthus species (Miscanthus sinensis and Miscanthus floridulus) of Taiwan and further explored the roles of host plant on root-associated microbiomes. Our results suggest that host plant genetic variation, edaphic factors, and site had effects on the root endophytic and rhizosphere soil microbial community compositions in both Miscanthus sinensis and Miscanthus floridulus, with a greater effect of plant genetic variation observed for the root endophytic communities. Host plant genetic variation also exerted a stronger effect on core prokaryotic communities than on non-core prokaryotic communities in each microhabitat of two Miscanthus species. From rhizosphere soils to root endophytes, prokaryotic co-occurrence network stability increased, but fungal co-occurrence network stability decreased. Furthermore, we found root endophytic microbial communities in two Miscanthus species were more strongly driven by deterministic processes rather than stochastic processes. Root-enriched prokaryotic OTUs belong to Gammaproteobacteria, Alphaproteobacteria, Betaproteobacteria, Sphingobacteriia, and [Saprospirae] both in two Miscanthus species, while prokaryotic taxa enriched in the rhizosphere soil are widely distributed among different phyla. We provide empirical evidence that host genetic variation plays important roles in root-associated microbiome in Miscanthus. The results of this study have implications for future bioenergy crop management by providing baseline data to inform translational research to harness the plant microbiome to sustainably increase agriculture productivity.

54 ENVIRONMENTAL SCIENCES↗

Genome‐wide association and genomic prediction for yield and component traits of Miscanthus sacchariflorus

Abstract Accelerating biomass improvement is a major goal of Miscanthus breeding. The development and implementation of genomic‐enabled breeding tools, like marker‐assisted selection (MAS) and genomic selection, has the potential to improve the efficiency of Miscanthus breeding. The present study conducted genome‐wide association (GWA) and genomic prediction of biomass yield and 14 yield‐components traits in Miscanthus sacchariflorus . We evaluated a diversity panel with 590 accessions of M. sacchariflorus grown across 4 years in one subtropical and three temperate locations and genotyped with 268,109 single‐nucleotide polymorphisms (SNPs). The GWA study identified a total of 835 significant SNPs and 674 candidate genes across all traits and locations. Of the significant SNPs identified, 280 were localized in mapped quantitative trait loci intervals and proximal to SNPs identified for similar traits in previously reported Miscanthus studies, providing additional support for the importance of these genomic regions for biomass yield. Our study gave insights into the genetic basis for yield‐component traits in M. sacchariflorus that may facilitate marker‐assisted breeding for biomass yield. Genomic prediction accuracy for the yield‐related traits ranged from 0.15 to 0.52 across all locations and genetic groups. Prediction accuracies within the six genetic groupings of M. sacchariflorus were limited due to low sample sizes. Nevertheless, the Korea/NE China/Russia ( N = 237) genetic group had the highest prediction accuracy of all genetic groups (ranging 0.26–0.71), suggesting that with adequate sample sizes, there is strong potential for genomic selection within the genetic groupings of M. sacchariflorus . This study indicated that MAS and genomic prediction will likely be beneficial for conducting population‐improvement of M. sacchariflorus .

09 BIOMASS FUELS↗

Biomass yield in a genetically diverse Miscanthus sacchariflorus germplasm panel phenotyped at five locations in Asia, North America, and Europe

Abstract Miscanthus is a high‐yielding bioenergy crop that is broadly adapted to temperate and tropical environments. Commercial cultivation of Miscanthus is predominantly limited to a single sterile triploid clone of Miscanthus × giganteus , a hybrid between Miscanthus sacchariflorus and M. sinensis. To expand the genetic base of M. × giganteus , the substantial diversity within its progenitor species should be used for cultivar improvement and diversification. Here, we phenotyped a diversity panel of 605 M. sacchariflorus from six previously described genetic groups and 27 M. × giganteus genotypes for dry biomass yield and 16 yield‐component traits, in field trials grown over 3 years at one subtropical location (Zhuji, China) and four temperate locations (Foulum, Denmark; Sapporo, Japan; Urbana, Illinois; and Chuncheon, South Korea). There was considerable diversity in yield and yield‐component traits among and within genetic groups of M. sacchariflorus , and across the five locations. Biomass yield of M. sacchariflorus ranged from 0.003 to 34.0 Mg ha −1 in year 3. Variation among the genetic groups was typically greater than within, so selection of genetic group should be an important first step for breeding with M. sacchariflorus . The Yangtze 2x genetic group (=ssp. lutarioriparius ) of M. sacchariflorus had the tallest and thickest culms at all locations tested. Notably, the Yangtze 2x genetic group's exceptional culm length and yield potential were driven primarily by a large number of nodes (>29 nodes culm −1 average over all locations), which was consistent with the especially late flowering of this group. The S Japan 4x, the N China/Korea/Russia 4x, and the N China 2x genetic groups were also promising genetic resources for biomass yield, culm length, and culm thickness, especially for temperate environments. Culm length was the best indicator of yield potential in M. sacchariflorus . These results will inform breeders' selection of M. sacchariflorus genotypes for population improvement and adaptation to target production environments.

59 BASIC BIOLOGICAL SCIENCES↗

Characterization of the Ghd8 Flowering Time Gene in a Mini-Core Collection of Miscanthus sinensis

The optimal flowering time for bioenergy crop Miscanthus is essential for environmental adaptability and biomass accumulation. However, little is known about how genes controlling flowering in other grasses contribute to flowering regulation in Miscanthus. Here, we report on the sequence characterization and gene expression of Miscanthus sinensisGhd8, a transcription factor encoding a HAP3/NF-YB DNA-binding domain, which has been identified as a major quantitative trait locus in rice, with pleiotropic effects on grain yield, heading date and plant height. In M. sinensis, we identified two homoeologous loci, MsiGhd8A located on chromosome 13 and MsiGhd8B on chromosome 7, with one on each of this paleo-allotetraploid species’ subgenomes. A total of 46 alleles and 28 predicted protein sequence types were identified in 12 wild-collected accessions. Several variants of MsiGhd8 showed a geographic and latitudinal distribution. Quantitative real-time PCR revealed that MsiGhd8 expressed under both long days and short days, and MsiGhd8B showed a significantly higher expression than MsiGhd8A. The comparison between flowering time and gene expression indicated that MsiGhd8B affected flowering time in response to day length for some accessions. This study provides insight into the conserved function of Ghd8 in the Poaceae, and is an important initial step in elucidating the flowering regulatory network of Miscanthus.

geographic distribution↗

Training Population Optimization for Genomic Selection in Miscanthus

Miscanthus is a perennial grass with potential for lignocellulosic ethanol production. To ensure its utility for this purpose, breeding efforts should focus on increasing genetic diversity of the nothospecies Miscanthus × giganteus (M×g) beyond the single clone used in many programs. Germplasm from the corresponding parental species M. sinensis (Msi) and M. sacchariflorus (Msa) could theoretically be used as training sets for genomic prediction of M×g clones with optimal genomic estimated breeding values for biofuel traits. To this end, we first showed that subpopulation structure makes a substantial contribution to the genomic selection (GS) prediction accuracies within a 538-member diversity panel of predominately Msi individuals and a 598-member diversity panels of Msa individuals. We then assessed the ability of these two diversity panels to train GS models that predict breeding values in an interspecific diploid 216-member M×g F2 panel. Low and negative prediction accuracies were observed when various subsets of the two diversity panels were used to train these GS models. To overcome the drawback of having only one interspecific M×g F2 panel available, we also evaluated prediction accuracies for traits simulated in 50 simulated interspecific M×g F2 panels derived from different sets of Msi and diploid Msa parents. The results revealed that genetic architectures with common causal mutations across Msi and Msa yielded the highest prediction accuracies. Ultimately, these results suggest that the ideal training set should contain the same causal mutations segregating within interspecific M×g populations, and thus efforts should be undertaken to ensure that individuals in the training and validation sets are as closely related as possible.

59 BASIC BIOLOGICAL SCIENCES↗