Search NASA⌕ Search

SEARCH · Search NASA

Results for “Genetic Code”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Characterisation and comparative analysis of mitochondrial genomes of false, yellow, black and blushing morels provide insights on their structure and evolution

Morchella species have considerable significance in terrestrial ecosystems, exhibiting a range of ecological lifestyles along the saprotrophism-to-symbiosis continuum. However, the mitochondrial genomes of these ascomycetous fungi have not been thoroughly studied, thereby impeding a comprehensive understanding of their genetic makeup and ecological role. In this study, we analysed the mitogenomes of 30 Morchellaceae species, including yellow, black, blushing and false morels. These mitogenomes are either circular or linear DNA molecules with lengths ranging from 217 to 565 kbp and GC content ranging from 38% to 48%. Fifteen core protein-coding genes, 28–37 tRNA genes and 3–8 rRNA genes were identified in these Morchellaceae mitogenomes. The gene order demonstrated a high level of conservation, with the cox1 gene consistently positioned adjacent to the rnS gene and cob gene flanked by apt genes. Some exceptions were observed, such as the rearrangement of atp6 and rps3 in Morchella importuna and the reversed order of atp6 and atp8 in certain morel mitogenomes. However, the arrangement of the tRNA genes remains conserved. We additionally investigated the distribution and phylogeny of homing endonuclease genes (HEGs) of the LAGLIDADG (LAGs) and GIY-YIG (GIYs) families. A total of 925 LAG and GIY sequences were detected, with individual species containing 19–48HEGs. These HEGs were primarily located in the cox1, cob, cox2 and nad5 introns and their presence and distribution displayed significant diversity amongst morel species. These elements significantly contribute to shaping their mitogenome diversity. Overall, this study provides novel insights into the phylogeny and evolution of the Morchellaceae.

59 BASIC BIOLOGICAL SCIENCES↗

From chromatin to crop: epigenetic innovations in bioenergy systems

Energy crops encompass a diverse array of plant species cultivated primarily as a source of biomass for energy generation and biofuel production. As such, they play a pivotal role in the transition to sustainable energy systems. However, their productivity is often limited by environmental stresses, nutrient availability, and the need for optimized yield. While traditional breeding and genetic engineering have driven improvements, challenges such as narrow genetic diversity, long development cycles, trait instability, and unexpected gene interactions remain. Epigenetics offers a largely untapped opportunity to overcome these constraints by regulating gene expression through mechanisms that are dynamic, finely tuned, and responsive to environmental and developmental cues. Epigenetic modifications including DNA methylation, histone post-translational changes, and small non-coding RNAs influence nearly all aspects of plant development and physiology, including traits central to bioenergy crops. While these mechanisms are well characterized in model species such as Arabidopsis thaliana, they remain underexplored in many purpose-grown energy crops. This review summarizes the current state of knowledge of epigenetic regulation in bioenergy species, explores how these mechanisms can be leveraged to enhance crop resilience and productivity, and identifies gaps in our understanding. By characterizing epigenetic mechanisms and harnessing epigenetic variation, we can expand the toolkit for developing resilient, high-yielding bioenergy crops to meet future environmental and energy demands.

09 BIOMASS FUELS↗

Induced protein expression in Leptospira spp. and its application to CRISPR/Cas9 mutant generation

Abstract Expanding the genetic toolkit for Leptospira spp. is a crucial step toward advancing our understanding of the biology and virulence of these atypical bacteria. Pathogenic Leptospiraare responsible for over 1 million human leptospirosis cases annually and significantly impact domestic animals. Bovine leptospirosis causes substantial financial losses due to abortion, stillbirths, and suboptimal reproductive performance. The advent of the CRISPR/Cas9 system has marked a turning point in genetic manipulation, with applications across multiple Leptospira species. However, incorporating controlled protein expression into existing genetic tools could further expand their utility. We developed and demonstrated the functionality of IPTG-inducible heterologous protein expression in Leptospira spp. This system was applied for regulated expression of dead Cas9 (dCas9) to generate knockdown mutants, and Cas9 to produce knockout mutants by inducing double-strand breaks (DSB) into desired targets. IPTG-induced dCas9 expression enabled validation of essential genes and non-coding RNAs. Additionally, IPTG-controlled Cas9 expression combined with a constitutive non-homologous end-joining (NHEJ) system allowed for successful recovery of knockout mutants, even in the absence of IPTG. These newly controlled protein expression systems will advance studies on the basic biology and virulence ofLeptospira, as well as facilitate knockout mutant generation for improved veterinary vaccines.

Science & Technology - Other Topics↗

How initial conditions-, structural-, and parameter-based model uncertainty interact and influence predictions in permafrost ecosystems: Modeling Archive

This dataset contains model output and input data, as well as source code examples for the Terrestrial Ecosystem Model with the Dynamic Vegetation Model and Dynamic Organic Soil (DVM-DOS-TEM) for the field sites Imnavait creek and the Bonanza creek Long Term Ecological Research Network (LTER). The data covers simulations from the last glacial maximum (LGM) until 2100 for a selection of paleo scenarios, setting the mean temperature of the LGM up to 10°C lower than pre-industrial conditions. The model structure was modulated to represent various model versions, and this dataset contains the relevant changes in the source code. The raw output data, the processed statistical data, the setup and processing scripts as well as parameter value distribution files from a parameter sensitivity analysis are included as well. Model outputs include active layer depth, organic soil carbon, soil layer depths, gross primary productivity (GPP) with and without nitrogen limitation, net primary productivity (NPP), soil liquid water content, heterotrophic, maintenance, and growth respiration, soil temperature, and vegetation carbon (*.nc files). The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research.Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.

54 ENVIRONMENTAL SCIENCES↗

Extreme Longitudinal Compression of Optimized Beams for MEV Ultrafast Electron Diffraction (Final Technical Report)

We worked out the design of a high repetition rate MeV energy ultrafast electron diffraction instrument based on the existing Cornell photoinjector, which can readily be applied to the presented findings. This example is a blueprint of other similarly arranged UED setups. Using particle tracking simulations in conjunction with multiobjective genetic algorithm optimization, we explored the smallest bunch lengths, emittance, and probe spot sizes achievable. As two limits, we defined stroboscopic conditions (with single electrons per pulse) and operation with 10 5 electrons per bunch which may be suitable for single-shot diffraction images. In the stroboscopic case, the flexibility provided by the many cavity bunching and acceleration allows for longitudinal phase space linearization without a higher harmonic field, providing sub-fs bunch lengths at the sample. Given low emittance photoemission conditions, these small bunch lengths can be maintained with probe transverse sizes at the single micron (1 μm) scale and below. In the case of 10 5 electrons per pulse, we simulated state-of-the-art 5D brightness conditions: rms bunch lengths of 10 fs with 3-nm normalized emittances, while permitting repetition rates as high as 1.3 GHz. We showed that in conjunction with collimating apertures, a novel focusing scheme achieves very high-quality emittance compensation for the central core of the beam composing 40% of particles, for a resulting beam size of 5 μm (rms). Finally, to aid in the design of new SRF-based ultrafast electron diffraction machines, we simulated the trade-off between the number of cavities used and achievable bunch length and emittance. In the longitudinal dimension, we made use of the fact that MeV UED requires much lower energy than the 15-MeV maxi mum energy of Cornell’s CBETA injector, and we may therefore use several of the SRF cavities for bunch length compression. In practice, we used a genetic optimization algorithm to choose the phases and amplitudes of the cavities appropriately for optimal bunching. In the zero space charge case, we found that bunching and acceleration are distributed across the six cavities in a way that produces a linearizing effect. And we showed that the ultimate bunch length can be limited by time-of-flight differences arising from transverse size and transverse momentum spread. The space charge code developed and used for this development is now permanent part of the Bmad accelerator simulation code and has already contributed to other developments, e.g., for the EIC electron cooler design.

43 PARTICLE ACCELERATORS↗

Adaptive gene loss in the common bean pan-genome during range expansion and domestication

The common bean ( Phaseolus vulgaris L.) is a crucial legume crop and an ideal evolutionary model to study adaptive diversity in wild and domesticated populations. Here, we present a common bean pan-genome based on five high-quality genomes and whole-genome reads representing 339 genotypes. It reveals ~234 Mb of additional sequences containing 6,905 protein-coding genes missing from the reference, constituting 49% of all presence/absence variants (PAVs). More non-synonymous mutations are found in PAVs than core genes, probably reflecting the lower effective population size of PAVs and fitness advantages due to the purging effect of gene loss. Our results suggest pan-genome shrinkage occurred during wild range expansion. Selection signatures provide evidence that partial or complete gene loss was a key adaptive genetic change in common bean populations with major implications for plant adaptation. The pan-genome is a valuable resource for food legume research and breeding for climate change mitigation and sustainable agriculture.

59 BASIC BIOLOGICAL SCIENCES↗

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

Multi-Objective Optimization of Uranium Target Assembly–3: A Comparison of Genetic and Traditional Methods

Commonly produced as a byproduct of uranium fission, 99 Mo is a key medical isotope that is in high demand in the United States. An international goal is to switch from medical isotope production technologies that require highly enriched uranium to medical isotope production technologies that require only low-enriched uranium. Niowave Inc. is contributing to this goal by developing an accelerator-driven subcritical assembly called the Uranium Target Assembly (UTA). This work compares the performance of Dakota’s Multi-Objective Genetic Algorithm (MOGA) against traditional sensitivity analysis in the neutronic optimization of the UTA-3 system. The design objectives are k-eigenvalue (k eff ) and natural uranium fission power, which are directly correlated with the amount of 99 Mo produced. Dakota:MOGA did not perform as well as human engineering ingenuity in optimization studies with high numbers of input parameters, such as fuel rod type selection and fuel rod placement. However, Dakota:MOGA did outperform traditional sensitivity analysis in optimization studies with fewer than 20 parameters and revealed the degree to which each parameter influences the optimal design space for k eff and natural uranium fission power (to a lesser extent). As the design model became more complex in the final stage of design, the computational resources required to calculate the design objective values in the Monte Carlo N-Particle transport code from selected input parameter combinations limited Dakota:MOGA’s performance, and, unfortunately, human intervention was required to discern the optimal design space. In conclusion, future work will attempt to reduce computational resource constraints by incorporating areduced-order neutronics model into the optimization cycle.

Accelerator-driven systems↗

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗

Exon disruptive variants in Populus trichocarpa associated with wood properties exhibit distinct gene expression patterns

Abstract Forest trees may harbor naturally occurring exon disruptive variants (DVs) in their gene sequences, which potentially impact important ecological and economic phenotypic traits. However, the abundance and molecular regulation of these variants remain largely unexplored. Here, 24,420 DVs were identified by screening 1014Populus trichocarpafull genomes. The identified DVs were predominantly heterozygous with allelic frequencies below 5% (only 26% of DVs had frequencies greater than 5%). Using common garden‐grown trees, DVs were assessed for gene expression variation in the developing xylem, revealing that their gene expression can be significantly altered, particularly for homozygous DVs (in the range of 27%–38% of cases depending on the studied common garden). DVs were further investigated for their correlations with 13 wood quality traits, revealing that, among the 148 discovered DV associations, 15 correlated with more than one wood property and six genes had more than one DV in their coding sequences associated with wood traits. Approximately one‐third of DVs correlated with wood property variation also showed significant gene expression variation, confirming their non‐spurious impact. These findings offer potential avenues for targeted introduction of homozygous mutations using tree biotechnology, and while the exact mechanisms by which DVs may directly influence wood formation remain to be unraveled, this study lays the groundwork for further investigation.

Genetics & Heredity↗

Introducing GPU Acceleration into the Python-Based Simulations of Chemistry Framework

We introduce the first version of GPU4P Y SCF, a module that provides GPU acceleration of methods in P Y SCF. As a core functionality, this provides a GPU implementation of two-electron repulsion integrals (ERIs) for contracted basis sets comprising up to g functions using the Rys quadrature. As an illustration of how this can accelerate a quantum chemistry workflow, we describe how to use the ERIs efficiently in the integral-direct Hartree–Fock build and nuclear gradient construction. Benchmark calculations show a significant speedup of 2 orders of magnitude with respect to the multithreaded CPU Hartree–Fock code of P Y SCF and the performance comparable to other open-source GPU-accelerated quantum chemical packages, including GAMESS and QUICK, on a single NVIDIA A100 GPU.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

Identification of functional non-coding variants associated with orofacial cleft

Oral facial cleft (OFC) comprises cleft lip with or without cleft palate (CL/P) or cleft palate only. Genome wide association studies (GWAS) of isolated OFC have identified common single nucleotide polymorphisms (SNPs) in many genomic loci where the presumed effector gene (for example, IRF6 in the 1q32 locus) is expressed in embryonic oral epithelium. To identify candidates for functional SNPs at eight such loci we conduct a massively parallel reporter assay in a fetal oral epithelial cell line, revealing SNPs with allele-specific effects on enhancer activity. We filter these SNPs against chromatin-mark evidence of enhancers and test a subset in traditional reporter assays, which support the candidacy of SNPs at loci containing FOXE1, IRF6, MAFB, TFAP2A, and TP63. For two SNPs near IRF6 and one near FOXE1, we engineer the genome of induced pluripotent stem cells, differentiate the cells into embryonic oral epithelium, and discover allele-specific effects on the levels of effector gene expression, and, in two cases, the binding affinity of transcription factors FOXE1 or ETS2. Conditional analyses of GWAS data suggest the two functional SNPs near IRF6 account for the majority of risk for CL/P at this locus. This study connects genetic variation associated with OFC to mechanisms of pathogenesis.

Kumari, Priyanka↗

Dataset_for_Conserved_macromolecular_architecture_of_Poplar_secondary_cell_walls_revealed_by_ssNMR_and_atomistic_modeling

This dataset contains solid-state 13C NMR data and atomistic molecular dynamics simulation files supporting the study of nanoscale secondary cell wall architecture across 13 genetically diverse Populus trichocarpa genotypes grown under uniform greenhouse conditions in 13C-enriched CO2 atmospheres (~89% 13C enrichment).The dataset contains two collections of solid-state 13C NMR data. (1) 200 MHz data (Bruker Avance III HD, 4 mm HX probe, 10 kHz MAS): raw Bruker TopSpin experiment folders and DMFIT-exported ascii spectra for selective and non-selective 1D 13C-13C spin diffusion experiments (3000 ms mixing) used to quantify inter-polymer spatial proximities, and short-mixing (1 ms) reference spectra used for polymeric abundance quantification by spectral deconvolution. (2) 600 MHz data (Bruker Avance III, 1.6 mm PhoenixNMR HXY probe, 30 kHz MAS): raw Bruker TopSpin experiment folders containing 2D CORD, 2D CP-INADEQUATE, and 13C/1H relaxation (T1, T1rho) experiments for all 13 genotypes, with processed Excel workbooks per experiment type. Molecular dynamics simulation code, coordinate files, and analysis scripts (NAMD/CHARMM/Python) for six atomistic cell wall models are included. Summarized ssNMR data are compiled into a single excel file and subjected to statistical analysis. Multivariate analysis code (PCA, Pearson correlation) and summary data are provided as excel worksheets and Jupyter notebooks (Python 3).

09 BIOMASS FUELS↗

Comparative genomics reveals the high diversity and adaptation strategies of Polaromonas from polar environments

Abstract Background Bacteria from the genus Polaromonas are dominant phylotypes found in a variety of low-temperature environments in polar regions. The diversity and biogeographic distribution of Polaromonas have been largely expanded on the basis of 16 S rRNA gene amplicon sequencing. However, the evolution and cold adaptation mechanisms of Polaromonas from polar regions are poorly understood at the genomic level. Results A total of 202 genomes of the genus Polaromonas were analyzed, and 121 different species were delineated on the basis of average nucleotide identity (ANI) and phylogenomic placements. Remarkably, 8 genomes recovered from polar environments clustered into a separate clade (‘polar group’ hereafter). The genome size, coding density and coding sequences (CDSs) of the polar group were significantly different from those of other nonpolar Polaromonas . Furthermore, the enrichment of genes involved in carbohydrate and peptide metabolism was evident in the polar group. In addition, genes encoding proteins related to betaine synthesis and transport were increased in the genomes from the polar group. Phylogenomic analysis revealed that two different evolutionary scenarios may explain the adaptation of Polaromonas to cold environments in polar regions. Conclusions The global distribution of the genus Polaromonas highlights its strong adaptability in both polar and nonpolar environments. Species delineation significantly expands our understanding of the diversity of the Polaromonas genus on a global scale. In this study, a polar-specific clade was found, which may represent a specific ecotype well adapted to polar environments. Collectively, genomic insight into the metabolic diversity, evolution and adaptation of the genus Polaromonas at the genome level provides a genetic basis for understanding the potential response mechanisms of Polaromonas to global warming in polar regions.

54 ENVIRONMENTAL SCIENCES↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗