Search NASASearch

SEARCH · Search NASA

Results for “Sequence Homology”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

An alternative pocket for binding the N‐degrons by the UBR1 and UBR2 ubiquitin E3 ligases

The UBR family of ubiquitin ligases binds to N-termini of their targets (known as N-degron) to induce their ubiquitination and degradation via a conserved domain known as UBR-box. UBR1 and UBR2 share the highest sequence homology among the family, and substantial structural studies were previously performed for substrate binding by the UBR-boxes of UBR1 and UBR2. Here, we describe a new pocket in the UBR-boxes of UBR1 and UBR2 for binding the second residues of N-degrons through determining five co-crystal structures of the UBR-boxes with various N-degron peptides. Together with binding affinities measured by fluorescence polarization, we show that the two highly homologous UBR-boxes can interact with the second residue of an N-degron differently. In addition, the UBR-boxes undergo different conformational changes when binding N-degrons. Furthermore, we demonstrate that the sidechain of the third amino acid of an N-degron has no contribution to binding the UBR-boxes. These findings represent a new conceptual advancement for the UBR E3 ligases and the new insights described here can be leveraged for developing their selective ligands for research and potential therapies.

N-end rule

Evolutionary trajectory of transcription factors and selection of targets for metabolic engineering

Transcription factors (TFs) provide potentially powerful tools for plant metabolic engineering as they often control multiple genes in a metabolic pathway. However, selecting the best TF for a particular pathway has been challenging, and the selection often relies significantly on phylogenetic relationships. Here, we offer examples where evolutionary relationships have facilitated the selection of the suitable TFs, alongside situations where such relationships are misleading from the perspective of metabolic engineering. We argue that the evolutionary trajectory of a particular TF might be a better indicator than protein sequence homology alone in helping decide the best targets for plant metabolic engineering efforts. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics

Genetics of Flooding Tolerance in an F 2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis Population

Miscanthus is a warm-season, perennial grass cultivated as a feedstock for bioenergy and bioproducts. M. sacchariflorus ssp. lutarioriparius has high yield potential and is well-adapted to seasonal flooding, but little is known about the genetics of this adaptation. We conducted a quantitative trait locus (QTL) analysis on a population of 332 diploid Miscanthus ×giganteus (Mxg) F2s derived from an initial cross between diploid M. sacchariflorus ssp. lutarioriparius ‘PF30022’ and diploid M. sinensis ‘PMS-014’, followed by intermating 50 F 1 s. Using tanks in a greenhouse to assess the effects of partial submergence on actively growing plants, we compared an aerobic soil control to a 6-week flood treatment. The study's primary objectives were to (1) identify QTL for flooding tolerance in Miscanthus , (2) identify candidate genes and (3) compare ethylene response factors in Miscanthus with those in rice and Arabidopsis , sorghum and maize for binding site sequence homology and synteny, especially those associated with flooding tolerance. In total, 10 QTL and 66 candidate genes for partial submergence tolerance were identified (including many for ethylene signalling), a first report for Miscanthus . Notably, none of the Miscanthus candidates were orthologs of rice Sub1A, SK1 or SK2 , yet the ‘PF30022’ parent exhibited a snorkeling phenotype, indicating convergent evolution. This study will facilitate breeding of climate-resiliant Mxg.

abiotic stress tolerance

Molecular and structural characterization of a Bacillus cereus strain producing an anthrax-like capsule

Bacillus cereus is a ubiquitous Gram-positive, spore-forming, rod-shaped saprophytic bacterium, occasionally reported to cause food-borne illnesses. However, instances of B. cereus strains harboring anthrax toxin and capsule genes have elevated certain strains as formidable pathogens and biothreats. This study focuses on the genomic analysis and the structural characterization of capsular material produced by the virulent B. cereus PATH2418 strain, isolated from the wound of a traumatic open fracture patient. The genome was sequenced using Nanopore MinION sequencing, revealing a chromosome of 5,270,283 bp and three plasmids. One plasmid, pATH1, was found to encode an operon for the biosynthesis of a bacterial capsule. This operon had sequence homology to the Bacillus anthracis capBCADE operon, which encodes the poly-γ-D-glutamate (PDGA) capsule. The capsule production in B. cereus PATH2418 was influenced by temperature and CO 2 levels. Structural analysis of the capsular material using a combined approach of nuclear magnetic resonance (NMR) and high-performance liquid chromatography (HPLC) techniques confirmed the presence of a high-molecular-weight poly-γ-glutamate capsule, with an enantiomeric composition of approximately 67% D-glutamic acid and 33% L-glutamic acid, matching that of B. anthracis.

Bacillus cereus

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES

Spatial proteomics reveals signal sequence characteristics correlated with localization in cyanobacteria

Abstract Cyanobacteria have an inner and outer cell membrane enclosing the periplasm and cell wall and an additional set of internal membranes (called the thylakoid membranes) enclosing the thylakoid lumen. The periplasm and thylakoid lumen have unique proteomes, but the mechanisms regulating protein sorting to these locations have remained elusive. Here, proximity-based proteomics using the engineered peroxidase APEX2 was performed in the cyanobacteria Synechococcus sp. PCC 7002 to profile the proteomes of the cytoplasm, thylakoid lumen, and the periplasm and outer membrane (P-OM). Our analyses revealed specific roles for the thylakoid lumen in photosynthesis and energy generation, as well as roles for the periplasm in metabolite transport and binding, cell motility, and cell wall maintenance. Forty proteins localized to both the thylakoid lumen and the P-OM; however, their biological functions remain unclear. We also analyzed the correlation between signal sequence characteristics and differential protein localization to either the thylakoid lumen or the P-OM. In PCC 7002, as well as Synechocystis sp. PCC 6803 and Nostoc sp. PCC 7120, thylakoid lumen proteins translocated across membranes via the Secretory (Sec) system possessed more hydrophobic and alpha-helical signal sequence H-regions than P-OM proteins. The signal sequences of homologous proteins in Gloeobacter violaceus PCC 7421, a cyanobacterial species with a combined thylakoid lumen and periplasmic space, did not exhibit such differences. Therefore, the pattern of increased H-region hydrophobicity and alpha helix content is specific to cyanobacteria with a separate thylakoid lumen space and likely contributes to proper protein sorting between the thylakoid lumen and periplasm.

Plant Sciences

Revisiting synthetic lethality of Gcn5-related N-acetyltransferase (GNAT) family mutations in Haloferax volcanii

ABSTRACT Lysine acetylation is a post-translational modification that occurs in all domains of life, highlighting its evolutionary significance. Previous genome comparison identified three Gcn5-related N-acetyltransferase (GNAT) family members as lysine acetyltransferase homologs (Pat1, Pat2, and Elp3) and two deacetylase homologs (Sir2 and HdaI) in the halophilic archaeonHaloferax volcanii, withelp3andpat2proposed as a synthetic lethal gene pair. Here, we advance these findings by performing single and double mutagenesis ofelp3with thepat1andpat2lysine acetyltransferase gene homologs. Genome sequencing and PCR screens of these strains reveal successful generation of Δelp3,Δpat1Δelp3, and Δpat2Δelp3mutant strains. Although these mutant strains exhibited a reduced growth rate compared to the parent, they remained viable. Overall, this study provides genetic evidence thatelp3andpat2, while impacting cell growth, are not a synthetic lethal gene pair as previously reported. IMPORTANCE Here, we reveal by whole-genome sequencing that the GNAT family gene homologselp3andpat2can be deleted in the sameHaloferax volcaniistrain. Beyond the targeted deletions, minimal differences between the parent and Δelp3Δpat2mutant were observed, suggesting that suppressor mutations are not responsible for our ability to generate this double mutant strain. Elp3 and Pat2, thus, may not share as close a functional relationship as implied by earlier study. Our finding is significant as Elp3 is thought to function in acetylation in tRNA modification, while Pat2 likely functions in the lysine acetylation of proteins.

Microbiology

Functional role of myosin-binding protein H in thick filaments of developing vertebrate fast-twitch skeletal muscle

Myosin-binding protein H (MyBP-H) is a component of the vertebrate skeletal muscle sarcomere with sequence and domain homology to myosin-binding protein C (MyBP-C). Whereas skeletal muscle isoforms of MyBP-C (fMyBP-C, sMyBP-C) modulate muscle contractility via interactions with actin thin filaments and myosin motors within the muscle sarcomere “C-zone,” MyBP-H has no known function. This is in part due to MyBP-H having limited expression in adult fast-twitch muscle and no known involvement in muscle disease. Quantitative proteomics reported here reveal that MyBP-H is highly expressed in prenatal rat fast-twitch muscles and larval zebrafish, suggesting a conserved role in muscle development and prompting studies to define its function. We take advantage of the genetic control of the zebrafish model and a combination of structural, functional, and biophysical techniques to interrogate the role of MyBP-H. Transgenic, FLAG-tagged MyBP-H or fMyBP-C both localize to the C-zones in larval myofibers, whereas genetic depletion of endogenous MyBP-H or fMyBP-C leads to increased accumulation of the other, suggesting competition for C-zone binding sites. Does MyBP-H modulate contractility in the C-zone? Globular domains critical to MyBP-C’s modulatory functions are absent from MyBP-H, suggesting that MyBP-H may be functionally silent. However, our results suggest an active role. In vitro motility experiments indicate MyBP-H shares MyBP-C’s capacity as a molecular “brake.” These results provide new insights and raise questions about the role of the C-zone during muscle development.

59 BASIC BIOLOGICAL SCIENCES

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES

Cofactor maturase NifEN: A prototype ancient nitrogenase?

Nitrogenase plays a key role in the global nitrogen cycle; yet, the evolutionary history of nitrogenase and, particularly, the sequence of appearance between the homologous, yet distinct NifDK (the catalytic component) and NifEN (the cofactor maturase) of the extant molybdenum nitrogenase, remains elusive. Here, we report the ability of NifEN to reduce N 2 at its surface-exposed L-cluster ([Fe 8 S 9 C]), a structural/functional homolog of the M-cluster (or cofactor; [(R-homocitrate)MoFe 7 S 9 C]) of NifDK. Furthermore, we demonstrate the ability of the L-cluster–bound NifDK to mimic its NifEN counterpart and enable N 2 reduction. These observations, coupled with phylogenetic, ecological, and mechanistic considerations, lead to the proposal of a NifEN-like, L-cluster–carrying protein as an ancient nitrogenase, the exploration of which could shed crucial light on the evolutionary origin of nitrogenase and related enzymes.

59 BASIC BIOLOGICAL SCIENCES

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Biophysical and biochemical evidence for the role of acetate kinases (AckAs) in an acetogenic pathway in pathogenic spirochetes

Unraveling the metabolism of Treponema pallidum is a key component to understanding the pathogenesis of the human disease that it causes, syphilis. For decades, it was assumed that glucose was the sole carbon/energy source for this parasitic spirochete. But the lack of citric-acid-cycle enzymes suggested that alternative sources could be utilized, especially in microaerophilic host environments where glycolysis should not be robust. Recent bioinformatic, biophysical, and biochemical evidence supports the existence of an acetogenic energy-conservation pathway in T . pallidum and related treponemal species. In this hypothetical pathway, exogenous D-lactate can be utilized by the bacterium as an alternative energy source. Herein, we examined the final enzyme in this pathway, acetate kinase (named TP0476), which ostensibly catalyzes the generation of ATP from ADP and acetyl-phosphate. We found that TP0476 was able to carry out this reaction, but the protein was not suitable for biophysical and structural characterization. We thus performed additional studies on the homologous enzyme (75% amino-acid sequence identity) from the oral pathogen Treponema vincentii , TV0924. This protein also exhibited acetate kinase activity, and it was amenable to structural and biophysical studies. We established that the enzyme exists as a dimer in solution, and then determined its crystal structure at a resolution of 1.36 Å, showing that the protein has a similar fold to other known acetate kinases. Mutation of residues in the putative active site drastically altered its enzymatic activity. A second crystal structure of TV0924 in the presence of AMP (at 1.3 Å resolution) provided insight into the binding of one of the enzyme’s substrates. On balance, this evidence strongly supported the roles of TP0476 and TV0924 as acetate kinases, reinforcing the hypothesis of an acetogenic pathway in pathogenic treponemes.

Deka, Ranjit K.

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES

Microbial species and intraspecies units exist and are maintained by ecological cohesiveness coupled to high homologous recombination

Abstract Recent genomic analyses have revealed that microbial communities are predominantly composed of persistent, sequence-discrete species and intraspecies units (genomovars), but the mechanisms that create and maintain these units remain unclear. By analyzing closely-related isolate genomes from the same or related samples and identifying recent recombination events using a novel bioinformatics methodology, we show that high ecological cohesiveness coupled to frequent-enough and unbiased (i.e., not selection-driven) horizontal gene flow, mediated by homologous recombination, often underlie these diversity patterns. Ecological cohesiveness was inferred based on greater similarity in temporal abundance patterns of genomes of the same vs. different units, and recombination was shown to affect all sizable segments of the genome (i.e., be genome-wide) and have two times or greater impact on sequence evolution than point mutations. These results were observed in bothSalinibacter ruber, an environmental halophilic organism, andEscherichia coli, the model gut-associated organism and an opportunistic pathogen, indicating that they may be more broadly applicable to the microbial world. Therefore, our results represent a departure compared to previous models of microbial speciation that invoke either ecology or recombination, but not necessarily their synergistic effect, and answer an important question for microbiology: what a species and a subspecies are.

Science & Technology - Other Topics