Search NASA⌕ Search

SEARCH · Search NASA

Results for “evolutionary”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Harnessing evolution: leveraging bacterial isoprenoid pathway diversity toward improved bioengineering strategies

Isoprenoids play vital roles in all domains of life, from beta-carotene in bacteria to heme in humans. Two distinct metabolic pathways have evolved to synthesize the critical precursor of all mature isoprenoids: the mevalonate (MEV) and the methylerythritol phosphate (MEP) pathways. Here, we quantify the extensive inter- and intra-genus heterogeneity in the usage of these two pathways with particular emphasis on rare bacteria that encode both, or neither, pathways. Furthermore, MEP intermediates themselves have non-isoprenogenic roles that may underlie evolutionary pressures driving pathway diversification. Understanding isoprenoid biosynthesis in bacteria offers new avenues toward more sustainable engineering of economically relevant molecules in microbes.

Biotechnology and Synthetic Biology↗

An RNA ligase partner for the prokaryotic protein-only RNase P: insights into the functional diversity of RNase P from genome mining

RNase P can use either an RNA- or a protein-based active site to catalyze 5'-maturation of transfer RNAs (tRNAs). This distinctive attribute in the biocatalytic repertoire raises questions about the underlying evolutionary driving forces, especially if each variant somehow affords a selective advantage under certain conditions. Upon mining all publicly available prokaryotic genomes and examining gene co-occurrence, we discovered that an RNA ligase with circularization activity was significantly overrepresented in genomes that contain the protein form of RNase P. This unexpected linkage inspires testable ideas to understand the bases for scenarios that might favor RNase P variants of different architectures/make-up.

HARP↗

Synthetic overlapping genes stabilize genetic systems

Overlapping genes—wherein two different proteins are translated from alternative reading frames of the same DNA sequence—provide a means to stabilize an engineered gene by directly linking its evolutionary fate with that of an overlapping gene. However, creating overlapping gene pairs is challenging, as it requires redesigning both protein products to accommodate overlap constraints. Here, we present a new “overlapping, alternate-frame insertion” (OAFI) method for creating synthetic overlapping genes by inserting an “inner” gene, encoded in an alternate frame, into a flexible region of an “outer” gene. Using OAFI, we create new overlapping gene pairs of genetic reporters and bacterial toxins within an antibiotic resistance gene. We show that both the inner and outer genes retain function despite redesign, with translation of the inner gene influenced by its overlap position in the outer gene. Importantly, we show that, despite these inner gene sequences not contributing to outer gene function, selection for the outer gene alters the permitted inactivating mutations in the inner gene, and that overlapping toxins can restrict horizontal gene transfer of the antibiotic resistance gene. Overall, OAFI offers a versatile tool for synthetic biology, expanding the applications of overlapping genes in gene stabilization and biocontainment.

Biological and medical sciences↗

Catabolic pathway acquisition by rhizosphere bacteria readily enables growth with a root exudate component but does not affect root colonization

Horizontal gene transfer (HGT) is a fundamental evolutionary process that plays a key role in bacterial evolution. The likelihood of a successful transfer event is expected to depend on the precise balance of costs and benefits resulting from pathway acquisition. Most experimental analyses of HGT have focused on phenotypes that have large fitness benefits under appropriate selective conditions, such as antibiotic resistance. However, many examples of HGT involve phenotypes that are predicted to provide smaller benefits, such as the ability to catabolize additional carbon sources. We have experimentally simulated the consequences of one such HGT event in the laboratory, studying the effects of transferring a pathway for catabolism of the plant-derived aromatic compound salicyl alcohol between rhizosphere isolates from the Pseudomonas genus. We find that pathway acquisition enables rapid catabolism of salicyl alcohol with only minor disruptions to the existing metabolic and regulatory networks of the new host. However, this new catabolic potential does not confer a measurable fitness advantage during competitive growth in the rhizosphere. We conclude that the phenotype of salicyl alcohol catabolism is readily transferable but is selectively neutral under environmentally relevant conditions. We propose that this condition is common and that HGT of many pathways will be self-limiting because the selective benefits are small.

59 BASIC BIOLOGICAL SCIENCES↗

Revisiting synthetic lethality of Gcn5-related N-acetyltransferase (GNAT) family mutations in Haloferax volcanii

ABSTRACT Lysine acetylation is a post-translational modification that occurs in all domains of life, highlighting its evolutionary significance. Previous genome comparison identified three Gcn5-related N-acetyltransferase (GNAT) family members as lysine acetyltransferase homologs (Pat1, Pat2, and Elp3) and two deacetylase homologs (Sir2 and HdaI) in the halophilic archaeonHaloferax volcanii, withelp3andpat2proposed as a synthetic lethal gene pair. Here, we advance these findings by performing single and double mutagenesis ofelp3with thepat1andpat2lysine acetyltransferase gene homologs. Genome sequencing and PCR screens of these strains reveal successful generation of Δelp3,Δpat1Δelp3, and Δpat2Δelp3mutant strains. Although these mutant strains exhibited a reduced growth rate compared to the parent, they remained viable. Overall, this study provides genetic evidence thatelp3andpat2, while impacting cell growth, are not a synthetic lethal gene pair as previously reported. IMPORTANCE Here, we reveal by whole-genome sequencing that the GNAT family gene homologselp3andpat2can be deleted in the sameHaloferax volcaniistrain. Beyond the targeted deletions, minimal differences between the parent and Δelp3Δpat2mutant were observed, suggesting that suppressor mutations are not responsible for our ability to generate this double mutant strain. Elp3 and Pat2, thus, may not share as close a functional relationship as implied by earlier study. Our finding is significant as Elp3 is thought to function in acetylation in tRNA modification, while Pat2 likely functions in the lysine acetylation of proteins.

Microbiology↗

A minimal SufB 2 C 2 complex functions as a [4Fe-4S] cluster scaffold in methanogenic archaea

Iron-sulfur clusters are essential cofactors in all domains of life, yet their biogenesis in obligately anaerobic archaea remains poorly understood. Here, we characterized the minimal two-protein SUF system in methanogenic archaea, composed solely of SufB and SufC. Using Methanococcus maripaludis as a model, we demonstrate that the SUF proteins from its native host form a stable SufB 2 C 2 heterotetramer that binds a [4Fe-4S] cluster via three conserved cysteines in SufC. Mutations of conserved cysteine and histidine residues of SufB do not impair cluster binding. The complex interacts with the SAM-containing methanogenesis marker protein 10 (MmpX), suggesting direct Fe-S cluster transfer from SufB 2 C 2 to target proteins. Mutational analysis of Methanothermococcus thermolithotrophicus proteins confirmed that SufC is the primary cluster-binding component, while SufB enhances ATPase and cluster transfer activities. Evolutionary comparisons suggest that this two-protein SUF system represents an ancestral form of Fe-S cluster biogenesis.

59 BASIC BIOLOGICAL SCIENCES↗

The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science

Modern scientific discovery increasingly requires coordinating distributed facilities and heterogeneous resources, forcing researchers to act as manual workflow coordinators rather than scientists. Advances in AI leading to AI agents show exciting new opportunities that can accelerate scientific discovery by providing intelligence as a component in the ecosystem. However, it is unclear how this new capability would materialize and integrate in the real world. To address this, we propose a conceptual framework where workflows evolve along two dimensions which are intelligence (from static to intelligent) and composition (from single to swarm) to chart an evolutionary path from current workflow management systems to fully autonomous scientific laboratories. With these trajectories in mind, we present an architectural blueprint that can help the community take the next steps towards harnessing the opportunities in autonomous science with the potential for 100x discovery acceleration and transformational scientific workflows.

Shin, Woong [ORNL] (ORCID:0000000172077814)↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Behavior, Energy, Autonomy, Mobility Modeling Framework (BEAM) v1.0

The Behavior, Energy, Autonomy, and Mobility (BEAM) model is an integrated, agent-based travel demand simulation framework. Individual agents express preferences through a utility- maximizing evolutionary algorithm that minimizes each individual’s cost and time spent traveling via diverse modal options, including the competition for scarce supply resources such as parking spaces and charging infrastructure. BEAM simulates the essential elements that compose a dynamic transportation system. From the road network, parking and charging infrastructure, to the transit system and a synthetic population with plans and preferences, the virtual system is an amalgamation of multiple spatially resolved layers that together represent an integrated transportation system. BEAM is an extension to the MATSim (Multi-Agent Transportation Simulation) model, where agents employ reinforcement learning across successive simulated days to maximize their personal utility through plan mutation (exploration) and selecting between previously executed plans (exploitation). The BEAM model shifts some of the behavioral emphasis in MATSim from across-day planning to within- day planning, where agents dynamically respond to the state of the system during the mobility simulation. In BEAM, agents can plan across all major modes of travel including driving, walking, biking, transit, and demand-responsive ride hailing. It is designed to integrate with other open source transportation models, such as ActivitySim.

Lazarus, Jessica↗

Poplar

SAND2025-00683O Poplar is a software tool that generates a phylogenetic tree from input gene and genome sequences. It integrates established tools to identify genes within genomes, group sequences, construct gene trees, and infer a species tree. Poplar processes nucleotide sequences, identifies similar sequences using Nucleotide BLAST, groups them with DBSCAN, aligns sequences with MAFFT, constructs gene trees with RAxML-NG, and infers a species tree using ASTRAL-Pro3. This pipeline provides a structured approach to phylogenetic analysis, facilitating the study of evolutionary relationships among species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krishnakumar, Raga↗

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]↗

Plant Reload Optimization (prlo)

The PRLO framework is built on a modular and extensible architecture that tightly couples advanced evolutionary optimization algorithms with nuclear fuel depletion solvers (i.e., nuclear physics neutronics code). It supports exploring complex, high-dimensional design spaces constrained by user-specified operational, safety, and economic constraints. Objectives such as minimizing fresh fuel enrichment, flattening radial and axial power distributions, and maximizing discharge burnup are evaluated. PRLO’s equilibrium cycle optimization capability enables the identification of core configurations that maintain fuel cycle sustainability over extended planning horizons. Its integration with the RAVEN platform facilitates optimization of loading patterns or fuel shuffling schemes across multiple cycles. The interface with SIMULATE, a licensed industry-standard nodal code developed by Studsvik, ensures accurate neutronic and thermal-hydraulic feedback for reactor core design. PRLO’s automated workflow engine supports iterative design refinement, enabling utilities to streamline core design processes and meet evolving performance and regulatory targets.

Kim, Junyung [Idaho National Laboratory] (00090005↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

LATTE_SciFM

LATTE: LAtent Token Transformer for Evolutionary dynamics

Most, Alex↗

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator↗

New species and records of the symbiotic shrimp genus Leptalpheus Williams, 1965, with notes on Fenneralpheus Felder & Manning, 1986, and preliminary molecular analysis of phylogenetic relationships (Crustacea: Decapoda: Alpheidae)

The shrimp genera Leptalpheus Williams, 1965 and Fenneralpheus Felder & Manning, 1986 are composed entirely of symbiotic species that co-inhabit burrows of infaunal macrocrustaceans. We report extensive collections of these genera from western Atlantic, eastern Pacific and Indo-West Pacific regions. Integrative taxonomy methods, including morphological comparisons and analysis of three mitochondrial genetic markers, are used to test species hypotheses and evolutionary relationships among members of these genera. Our molecular analysis failed to recover Leptalpheus or Fenneralpheus as monophyletic groups. Our results strongly supported the monophyly of three clades composed of species of Leptalpheus, loosely corresponding to previously proposed species groups. Three new species closely related to Leptalpheus forceps Williams, 1965, L. marginalis Anker, 2011, and L. mexicanus Ríos & Carvacho, 1983 are described. Leptalpheus ankeri n. sp., from the Caribbean Sea, Atlantic coast of Florida, and Gulf of Mexico, is a polymorphic species that exhibits two major cheliped morphotypes. Leptalpheus sibo n. sp., from the Pacific coast of Nicaragua, is morphologically very similar to L. ankeri n. sp., likely its transisthmian sister species, and shares its cheliped polymorphism. A reassessment of L. forceps concluded that records of this species from the Caribbean Sea and Brazil are not conspecific with L. forceps sensu stricto from the Atlantic coast of the USA and the Gulf of Mexico, and they are herein described as Leptalpheus degravei n. sp. Based on both molecular and morphological evidence, we found Leptalpheus bicristatus Anker, 2011 to be a junior synonym of L. mexicanus and Leptalpheus canterakintzi Anker & Lazarus, 2015 to be a junior synonym of Leptalpheus azuero Anker, 2011. First reports of Leptalpheus axianassae Dworschak & Coelho, 1999 in Texas and Mexico, Leptalpheus denticulatus Anker & Marin, 2009 in the Mariana Islands, Leptalpheus felderi Anker, Vera Caripe & Lira, 2006 and Leptalpheus lirai Vera Caripe, Pereda & Anker, 2021 in the USA, and Leptalpheus pereirai Anker & Vera Caripe, 2016 in Cuba are included.

Zoology↗

Chromosome-level genome assemblies and genetic maps reveal heterochiasmy and macrosynteny in endangered Atlantic Acropora

Abstract Background Over their evolutionary history, corals have adapted to sea level rise and increasing ocean temperatures, however, it is unclear how quickly they may respond to rapid change. Genome structure and genetic diversity contained within may highlight their adaptive potential. Results We present chromosome-scale genome assemblies and linkage maps of the critically endangered Atlantic acroporids,Acropora palmataandA. cervicornis. Both assemblies and linkage maps were resolved into 14 chromosomes with their gene content and colinearity. Repeats and chromosome arrangements were largely preserved between the species. The family Acroporidae and the genusAcroporaexhibited many phylogenetically significant gene family expansions. Macrosynteny decreased with phylogenetic distance. Nevertheless, scleractinians shared six of the 21 cnidarian ancestral linkage groups as well as numerous fission and fusion events compared to other distantly related cnidarians. Genetic linkage maps were constructed from oneA. palmatafamily and 16A. cervicornisfamilies using a genotyping array. The consensus maps span 1,013.42 cM and 927.36 cM forA. palmataandA. cervicornis, respectively. Both species exhibited high genome-wide recombination rates (3.04 to 3.53 cM/Mb) and pronounced sex-based differences, known as heterochiasmy, with 2 to 2.5X higher recombination rates estimated in the female maps. Conclusions Together, the chromosome-scale assemblies and genetic maps we present here are the first detailed look at the genomic landscapes of the critically endangered Atlantic acroporids. These data sets revealed that adaptive capacity of Atlantic acroporids is not limited by their recombination rates. The sister species maintain macrosynteny with few genes with high sequence divergence that may act as reproductive barriers between them. In the AtlanticAcropora, hybridization between the two sister species yields an F1 hybrid with limited fertility despite the high levels of macrosynteny and gene colinearity of their genomes. Together, these resources now enable genome-wide association studies and discovery of quantitative trait loci, two tools that can aid in the conservation of these species.

Biotechnology & Applied Microbiology↗

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES↗