Search NASA⌕ Search

DOE OSTI · 1854526

Predicting Antimicrobial Resistance Using Partial Genome Alignments

Abstract

Antimicrobial resistance (AMR) is an important global health threat that impacts millions of people worldwide each year. Developing methods that can detect and predict AMR phenotypes can help to mitigate the spread of AMR by informing clinical decision making and appropriate mitigation strategies. Many bioinformatic methods have been developed for predicting AMR phenotypes from whole-genome sequences and AMR genes, but recent studies have indicated that predictions can be made from incomplete genome sequence data. In order to more systematically understand this, we built random forest-based machine learning classifiers for predicting susceptible and resistant phenotypes for Klebsiella pneumoniae (1,640 strains), Mycobacterium tuberculosis (2,497 strains), and Salmonella enterica (1,981 strains). We started by building models from alignments that were based on a reference chromosome for each species. We then subsampled each chromosomal alignment and built models for the resulting subalignments, finding that very small regions, representing approximately 0.1 to 0.2% of the chromosome, are predictive. In K. pneumoniae, M. tuberculosis, and S. enterica, the subalignments are able to predict multiple AMR phenotypes with at least 70% accuracy, even though most do not encode an AMR-related function. We used these models to identify regions of the chromosome with high and low predictive signals. Finally, subalignments that retain high accuracy across larger phylogenetic distances were examined in greater detail, revealing genes and intergenic regions with potential links to AMR, virulence, transport, and survival under stress conditions. IMPORTANCE Antimicrobial resistance causes thousands of deaths annually worldwide. Understanding the regions of the genome that are involved in antimicrobial resistance is important for developing mitigation strategies and preventing transmission. Machine learning models are capable of predicting antimicrobial resistance phenotypes from bacterial genome sequence data by identifying resistance genes, mutations, and other correlated features. They are also capable of implicating regions of the genome that have not been previously characterized as being involved in resistance. In this study, we generated global chromosomal alignments for Klebsiella pneumoniae, Mycobacterium tuberculosis, and Salmonella enterica and systematically searched them for small conserved regions of the genome that enable the prediction of antimicrobial resistance phenotypes. In addition to known antimicrobial resistance genes, this analysis identified genes involved in virulence and transport functions, as well as many genes with no previous implication in antimicrobial resistance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aytan-Aktug, D., Nguyen, M., Clausen, P. C., Stevens, R. L., Aarestrup, F. M., Lund, O., Davis, J. J.. 2021-06-15. Predicting Antimicrobial Resistance Using Partial Genome Alignments. https://doi.org/10.1128/msystems.00185-21

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-resolved insights into microbial diversity and elemental cycling in Winogradsky columns

We retained 18 MAGs with ≥50% completion and <10% contamination (i.e., at least medium quality). Of these, 10 had >90% completion and <5% contamination; however, only one (Paceibacteria Bin.003_MG) can be described as high-quality, as the others lacked a full suite of 5S, 16S, and 23S rRNA genes. To maximize the diversity of our recovered MAGs, we also retained one MAG (Chromatiaceae Bin.008_AM) with >40% (but less than 50%) completion and <5% contamination, as well as one (Rhodopseudomonas Bin.015_MK) with >90% completion and <20% (but>10%) contamination. Interestingly, significant chimerism was not detected in this MAG (40) , suggesting that the elevated contamination (20%) may instead reflect two closely related strains collapsing into a single bin. Consistent with this, contig coverage was bimodal, with roughly 17% of the assembly at ~115x and the remaining 83% at ~282x, while GC content remained uniform across both groups (~64%), arguing against contamination from a taxonomically distinct source.

59 BASIC BIOLOGICAL SCIENCES↗