Search NASA⌕ Search

DOE OSTI · 3676828

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Abstract

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kapoor, Muskan, Ventura, Enrique, Walsh, Amy, Sokolov, Alexey, George, Nancy, Kumari, Sunita, Provart, Nicholas, Cole, Benjamin, Libault, Marc, Tickle, Timothy, Warren, Wesley, Koltes, James, Papatheodorou, Irene, Ware, Doreen, Harrison, Peter, Elsik, Christine, Yordanova, Galabina, Burdett, Tony, Tuggle, Christopher. 2024-11-29. Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research. https://doi.org/10.3389/fgene.2024.1460351

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Development and characterization of a wild emmer wheat backcross introgression population for hard winter wheat improvement

Abstract Wild emmer wheat (Triticum turgidumsubsp.dicoccoides) is the tetraploid progenitor of hexaploid bread wheat (Triticum aestivumL.) and is known to be a valuable source of genetic variation for wheat improvement. However, direct evaluation of wild emmer diversity for agronomic potential has limited value unless performed in the backgrounds of adapted cultivars. Here, we present a genetic characterization of a population of 1601 backcross recombinant inbred lines, with an average genome composition of 75% bread wheat and 25% wild emmer. Low‐coverage whole‐genome sequencing allowed introgressions and aneuploidies to be identified at a relatively low cost per sample. We identified a relatively large proportion of small introgressions (median length 38 Mb), and we found introgressions to be distributed across all chromosomes. Approximately 44% of genotyped progeny carried at least one aneuploidy, with monosomies being by far the most common. This population, which we have denoted as the Great Plains Wild Emmer/Hard Winter Wheat introgression population (GPWEW‐IP), is, to our knowledge, the first introgression population developed through the direct hybridization of wild emmer wheat and US‐adapted hard winter wheat. We believe that this population represents a valuable resource for wheat breeders and will accelerate the discovery and integration of useful variation from wild emmer wheat.

Genetics & Heredity↗

Phenome‐to‐genome insights for evaluating root system architecture in field studies of maize

Abstract Understanding the genetic basis of root system architecture (RSA) in crops requires innovative approaches that enable both high‐throughput and precise phenotyping in field conditions. In this study, we evaluated multiple phenotyping and analytical frameworks for quantifying RSA in mature, field‐grown maize in three field experiments. We used forward and reverse genetic approaches to evaluate >1700 maize root crowns, including a diversity panel, a biparental mapping population, and maize mutant and wild‐type alleles at two known RSA genes,DEEPER ROOTING 1(DRO1) andRootless1(Rt1). We show the utility of increasing the dimensionality of traditional two‐dimensional (2D) techniques, referred to as the “2D multi‐view” method, to improve the capture of whole root system information for mapping genetic variation influencing RSA. Comparison of univariate and multivariate genome‐wide association study (GWAS) approaches revealed that multivariate traits were effective at dissecting complex RSA phenotypes and identifying pleiotropic quantitative trait loci (QTLs). Overall, three‐dimensional (3D) root models generated from X‐ray computed tomography and digital phenotyping captured a larger proportion of RSA trait variations compared to other methods of root phenotyping, as evidenced by both genome‐wide and single‐gene analyses. Among the individual root traits, root pulling force emerged as a highly heritable estimate of RSA that identified the largest number of shared QTLs with 3D phenotypes. Our study shows that integrating complementary phenotyping technologies helps to provide a more comprehensive understanding of the genetic architecture of RSA in field‐grown maize.

Genetics & Heredity↗

Genetic variation at transcription factor binding sites largely explains phenotypic heritability in maize

Abstract Comprehensive maps of functional variation at transcription factor (TF) binding sites (cis-elements) are crucial for elucidating how genotype shapes phenotype. Here, we report the construction of a pan-cistrome of the maize leaf under well-watered and drought conditions. We quantified haplotype-specific TF footprints across a pan-genome of 25 maize hybrids and mapped over 200,000 variants, genetic, epigenetic, or both (termed binding quantitative trait loci (bQTL)), linked tocis-element occupancy. Three lines of evidence support the functional significance of bQTL: (1) coincidence with causative loci that regulate traits, includingvgt1,ZmTRE1and the MITE transposon nearZmNAC111under drought; (2) bQTL allelic bias is shared between inbred parents and matches chromatin immunoprecipitation sequencing results; and (3) partitioning genetic variation across genomic regions demonstrates that bQTL capture the majority of heritable trait variation across ~72% of 143 phenotypes. Our study provides an auspicious approach to make functionalcis-variation accessible at scale for genetic studies and targeted engineering of complex traits.

Genetics & Heredity↗