Search NASA⌕ Search

SEARCH · Search NASA

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Parallel measurement of transcriptomes and proteomes from same single cells using nanodroplet splitting

Single-cell multiomics provides comprehensive insights into gene regulatory networks, cellular diversity, and temporal dynamics. Here, we introduce nanoSPLITS (nanodroplet SPlitting for Linked-multimodal Investigations of Trace Samples), an integrated platform that enables global profiling of the transcriptome and proteome from same single cells via RNA sequencing and mass spectrometry-based proteomics, respectively. Benchmarking of nanoSPLITS demonstrates high measurement precision with deep proteomic and transcriptomic profiling of single-cells. We apply nanoSPLITS to cyclin-dependent kinase 1 inhibited cells and found phospho-signaling events could be quantified alongside global protein and mRNA measurements, providing insights into cell cycle regulation. We extend nanoSPLITS to primary cells isolated from human pancreatic islets, introducing an efficient approach for facile identification of unknown cell types and their protein markers by mapping transcriptomic data to existing large-scale single-cell RNA sequencing reference databases. Accordingly, we establish nanoSPLITS as a multiomic technology incorporating global proteomics and anticipate the approach will be critical to furthering our understanding of biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

Growth-induced Donnan exclusion influences swelling kinetics in highly charged dynamic polymerization hydrogels

Polymeric gels crosslinked by DNA sequences can exploit DNA strand-displacement reactions to promote swelling through dynamic polymerization. The degree of swelling and the rate of swelling must be directly tunable to achieve the promise of programmable soft matter. Though the kinetics of the strand-displacement reaction provide insertion rates up to 10 4 /Molar/second as measured in bulk solution, DNA hydrogel swelling can take upwards of 30 h to complete. Computational modeling of the reaction-induced swelling of these gels with our recently-developed reactive electrochemomechanical theory (Zimmerman et al., 2024) suggests that their extraordinarily slow swelling is partly due to a scaling mismatch between the addition of charge and the addition of fluid volume, leading to a large transient increase in the fixed charge density. The significant increase in the gel’s fixed charge density, due to the binding of negatively charged DNA, sharply restricts the concentration of mobile hairpins through the phenomenon of Donnan charge exclusion, an effect commonly exploited in nanofiltration applications using polymeric membranes. The scaling problem is overcome when the mean additional swelling provided to the hydrogel by addition of a crosslink is above a critical value, thus the swelling outpaces the charge accumulation, leading the fixed charge density to drop and significantly accelerating the swelling process. This study shows that Donnan exclusion can explain the kinetics of DNA hydrogel swelling, and studies ways to modulate the reaction speed by either modifying the salt concentration or increasing or decreasing the number of base pairs in each DNA sequence.

42 ENGINEERING↗

RT-EZ: A Golden Gate Assembly Toolkit for Streamlined Genetic Engineering of Rhodotorula toruloides

For economic and sustainable biomanufacturing, the oleaginous yeast Rhodotorula toruloides has emerged as a promising platform for producing biofuels, pharmaceuticals, and other valuable chemicals. However, genetic manipulation of R. toruloides has been limited by its high GC content and the lack of a replicating plasmid, necessitating gene integration into the genome of the yeast. To address these challenges, we developed the RT-EZ (R. toruloides Efficient Zipper) toolkit, a versatile tool based on Golden Gate assembly, designed to streamline R. toruloides engineering with improved efficiency and flexibility. The RT-EZ toolkit simplifies vector construction by incorporating new features such as bidirectional promoters and 2A peptides, color-based screening using RFP, and sequences optimized for both Agrobacterium tumefaciens-mediated transformation (ATMT) and easy linearization, enabling straightforward selection and transformation. Notably, the RT-EZ kit can be used to construct an expression cassette with four different genes in one assembly reaction, significantly improving vector construction speed and efficiency. The utility of the RT-EZ toolkit was demonstrated through the successful synthesis of arachidonic acid in R. toruloides by coexpressing fatty acid elongases and desaturases. Furthermore, this result underscores the potential of the RT-EZ toolkit to advance synthetic biology in R. toruloides, providing a streamlined method for addressing genetic engineering challenges in the yeast.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Testing of a Line Driver With Configurable Pre-Emphasis on Lossy Transmission Lines

Rare-event physics experiments such as the Deep Underground Neutrino Experiment (DUNE) or the next Enriched Xenon Observatory (nEXO) experiment search for rare, low-energy events, detected by sensitive detectors immersed in a cryogenic noble liquid (e.g., liquid argon or xenon). Readout electronics used within such detectors must consume minimal power while operating reliably in cryogenic environments. Furthermore, in the case of nEXO, maximizing the radiopurity of the environment is vital to minimize background noise, thus placing strict limits on the volume of dielectric materials, leading to high-loss data cables spanning distances up to 12 m. Such cables cause high attenuation and intersymbol interference (ISI), resulting in a high bit-error rate (BER). These issues were addressed by developing an integrated line driver with configurable pre-emphasis in a 65-nm CMOS process. The pre-emphasis parameters can be programmed to minimize BER for specific cables and data rates under power constraints. Here, the driver was tested at both room and cryogenic temperatures. In both cases, the output BER was found to be strongly correlated with the pre-emphasis settings. Furthermore, analysis and simulation showed that adapting the pre-emphasis settings based on the incoming bit sequence can further improve performance with minimal changes to the current solution.

47 OTHER INSTRUMENTATION↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "RT-EZ: A Golden Gate Assembly Toolkit for Streamlined Genetic Engineering of Rhodotorula toruloides"

For economic and sustainable biomanufacturing, the oleaginous yeast Rhodotorula toruloides has emerged as a promising platform for producing biofuels, pharmaceuticals, and other valuable chemicals. However, genetic manipulation of R. toruloides has been limited by its high GC content and the lack of a replicating plasmid, necessitating gene integration into the genome of the yeast. To address these challenges, we developed the RT-EZ ( R. toruloides Efficient Zipper) toolkit, a versatile tool based on Golden Gate assembly, designed to streamline R. toruloides engineering with improved efficiency and flexibility. The RT-EZ toolkit simplifies vector construction by incorporating new features such as bidirectional promoters and 2A peptides, color-based screening using RFP, and sequences optimized for both Agrobacterium tumefaciens-mediated transformation (ATMT) and easy linearization, enabling straightforward selection and transformation. Notably, the RT-EZ kit can be used to construct an expression cassette with four different genes in one assembly reaction, significantly improving vector construction speed and efficiency. The utility of the RT-EZ toolkit was demonstrated through the successful synthesis of arachidonic acid in R. toruloides by coexpressing fatty acid elongases and desaturases. This result underscores the potential of the RT-EZ toolkit to advance synthetic biology in R. toruloides , providing a streamlined method for addressing genetic engineering challenges in the yeast.

gene editing↗

Virulence and Genetic Diversity of Puccinia spp., Causal Agents of Rust on Switchgrass (Panicum virgatum L.) in the USA

Switchgrass (Panicum virgatum L.) is an important cellulosic biofuel grass native to North America. Rust, caused by Puccinia spp. is the most predominant disease of switchgrass and has the potential to impact biomass conversion. In this study, virulence patterns were determined on a set of 38 switchgrass genotypes for 14 single-spore rust isolates from 14 field samples collected in seven states. Single nucleotide polymorphism (SNP) variation was also assessed in 720 sequenced cloned amplicons representing 654 base pairs of the elongation factor 1-α gene from the field samples. Five major haplotypes were identified differing by 11 out of the 39 SNP positions identified. STRUCTURE, Principal Coordinate Analysis, and phylogenetic analyses divided the rust population into two genetic clusters. Virginia and Georgia had the highest and lowest rust genetic diversity, respectively. Only nine accessions showed a differential disease response between the 14 isolates, allowing the identification of eight races, differing by 1–3 virulence factors. Overall, the results suggested clonal reproduction of the pathogen and a North–South differentiation via local adaptation. However, similar haplotypes and races were also recovered from several states, suggesting migration events, and highlighting the need to further investigate the switchgrass rust population structure and evolution in the USA.

Bahri, Bochra A. (ORCID:0000000159055880)↗

In vivo mapping of mutagenesis sensitivity of human enhancers

Distant-acting enhancers are central to human development1. However, our limited understanding of their functional sequence features prevents the interpretation of enhancer mutations in disease2. Here we determined the functional sensitivity to mutagenesis of human developmental enhancers in vivo. Focusing on seven enhancers that are active in the developing brain, heart, limb and face, we created over 1,700 transgenic mice for over 260 mutagenized enhancer alleles. Systematic mutation of 12-base-pair blocks collectively altered each sequence feature in each enhancer at least once. We show that 69% of all blocks are required for normal in vivo activity, with mutations more commonly resulting in loss (60%) than in gain (9%) of function. Using predictive modelling, we annotated critical nucleotides at the base-pair resolution. The vast majority of motifs predicted by these machine learning models (88%) coincided with changes in in vivo function, and the models showed considerable sensitivity, identifying 59% of all functional blocks. Taken together, our results reveal that human enhancers contain a high density of sequence features that are required for their normal in vivo function and provide a rich resource for further exploration of human enhancer logic.

Kosicki, Michael↗

Microreactor Assembly Transportation Cask Model Description for Criticality Safety Validation Basis Assessment

Criticality safety analyses are completed on a transportation cask used for microreactor assembly shipment to provide an example of model and analysis to industry for reproducing this type of study on their microreactor fuel shipment. The fuel assembly considered is based on a gas-cooled microreactor (GC-MR), which utilizes HALEU fuel in the form of TRISO particles and utilizes various design options considered in industry designs. Various versions of this GC-MR assembly were studied, with and without YH2 moderator, providing similar conclusions. The shipment cask design is revised based on an existing design ES-3100, developed by Y-12 for the transport of highly enriched uranium (HEU), but is enlarged to hold the GC-MR fuel assembly. Criticality safety analysis for the cask/GC-MR fuel assembly package was performed using the CSAS6 sequence of SCALE6.3.2, utilizing the ENDF/B-VII.1 based continuous energy neutron library, and the analysis strictly follows the guideline from NRC reference reports. Different scenarios, e.g. normal operation, undamaged cask with water flooded, damaged cask with optimal water moderation, have been analyzed and it could be concluded the package would always have a large margin of subcriticality even packed in an infinite array. Sensitivity and similarity analyses are also performed using the TSUNAMI sequence of SCALE6.3.2, and the similarity analysis uses all the experiments from the ICSBEP Handbook with Intermediate and Mixed Enriched Uranium (IEU) and Low Enriched Uranium (LEU) systems together with additional ones that are sponsored by the DNCSH program. These similarity analyses indicate that dry cases have no similar benchmark experiments (ck values greater than 0.8), which may become problematic if more assemblies are shipped together (or a fully loaded core is shipped) and margin to criticality is reduced. However, the damaged cask models with flooded assemblies exhibited similarities to many experiments with ck values greater than 0.8.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Microreactor Assembly Transportation Cask Model Description for Criticality Safety Validation Basis Assessment (Rev. 3)

Criticality safety analyses are completed on a transportation cask used for microreactor assembly shipment to provide an example of model and analysis to industry for reproducing this type of study on their microreactor fuel shipment. The fuel assembly considered is based on a gas-cooled microreactor (GC-MR), which utilizes HALEU fuel in the form of TRISO particles and utilizes various design options considered in industry designs. Various versions of this GC-MR assembly were studied, with and without YH 2 moderator, providing similar conclusions. The shipment cask design is revised based on an existing ES-3100 design, developed by Y-12 for the transport of highly enriched uranium (HEU), but is enlarged to hold the GC-MR fuel assembly. Criticality safety analysis for the cask/GC-MR fuel assembly package was performed using the CSAS6 sequence of SCALE6.3.2, utilizing the ENDF/B-VII.1 based continuous energy neutron library, and the analysis strictly follows the guideline from NRC reference reports. Different scenarios, e.g. normal operation, undamaged cask with water flooded, damaged cask with optimal water moderation, have been analyzed and it could be concluded the package would always have a large margin of subcriticality even packed in an infinite array. Sensitivity and similarity analyses are also performed using the TSUNAMI sequence of SCALE6.3.2, and the similarity analysis uses all the experiments from the ICSBEP Handbook with Highly Enriched Uranium (HEU), Intermediate and Mixed Enriched Uranium (IEU) and Low Enriched Uranium (LEU) systems, together with additional ones that are sponsored by the DNCSH program, and selected IRPhEP experiments using TRISO fuel and graphite moderator. These similarity analyses indicate that the dry nominal design has no similar benchmark experiments (ck values greater than 0.8), which may become problematic if more assemblies are shipped together (or a fully loaded core is shipped) and margin to criticality is reduced. However, the cask models with flooded assemblies exhibited similarities to many experiments with c k values greater than 0.8.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning Framework for Conotoxin Class and Molecular Target Prediction

Conotoxins are small and highly potent neurotoxic peptides derived from the venom of marine cone snails which have captured the interest of the scientific community due to their pharmacological potential. These toxins display significant sequence and structure diversity, which results in a wide range of specificities for several different ion channels and receptors. Despite the recognized importance of these compounds, our ability to determine their binding targets and toxicities remains a significant challenge. Predicting the target receptors of conotoxins, based solely on their amino acid sequence, remains a challenge due to the intricate relationships between structure, function, target specificity, and the significant conformational heterogeneity observed in conotoxins with the same primary sequence. We have previously demonstrated that the inclusion of post-translational modifications, collisional cross sections values, and other structural features, when added to the standard primary sequence features, improves the prediction accuracy of conotoxins against non-toxic and other toxic peptides across varied datasets and several different commonly used machine learning classifiers. Here, we present the effects of these features on conotoxin class and molecular target predictions, in particular, predicting conotoxins that bind to nicotinic acetylcholine receptors (nAChRs). We also demonstrate the use of the Synthetic Minority Oversampling Technique (SMOTE)-Tomek in balancing the datasets while simultaneously making the different classes more distinct by reducing the number of ambiguous samples which nearly overlap between the classes. In predicting the alpha, mu, and omega conotoxin classes, the SMOTE-Tomek PCA PLR model, using the combination of the SS and P feature sets establishes the best performance with an overall accuracy (OA) of 95.95%, with an average accuracy (AA) of 93.04%, and an f1 score of 0.959. Using this model, we obtained sensitivities of 98.98%, 89.66%, and 90.48% when predicting alpha, mu, and omega conotoxin classes, respectively. Similarly, in predicting conotoxins that bind to nAChRs, the SMOTE-Tomek PCA SVM model, which used the collisional cross sections (CCSs) and the P feature sets, demonstrated the highest performance with 91.3% OA, 91.32% AA, and an f1 score of 0.9131. The sensitivity when predicting conotoxins that bind to nAChRs is 91.46% with a 91.18% sensitivity when predicting conotoxins that do not bind to nAChRs.

59 BASIC BIOLOGICAL SCIENCES↗

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING↗

Full-Length ASFV B646L Gene Sequencing by Nanopore Offers a Simple and Rapid Approach for Identifying ASFV Genotypes

African swine fever (ASF) is an acute, highly hemorrhagic viral disease in domestic pigs and wild boars. The disease is caused by African swine fever virus, a double stranded DNA virus of the Asfarviridae family. ASF can be classified into 25 different genotypes, based on a 478 bp fragment corresponding to the C-terminal sequence of the B646L gene, which is highly conserved among strains and encodes the major capsid protein p72. The C-terminal end of p72 has been used as a PCR target for quick diagnosis of ASF, and its characterization remains the first approach for epidemiological tracking and identification of the origin of ASF in outbreak investigations. Recently, a new classification of ASF, based on the complete sequence of p72, reduced the 25 genotypes into only six genotypes; therefore, it is necessary to have the capability to sequence the full-length B646L gene (p72) in a rapid manner for quick genotype characterization. Here, we evaluate the use of an amplicon approach targeting the whole B646L gene, coupled with nanopore sequencing in a multiplex format using Flongle flow cells, as an easy, low cost, and rapid method for the characterization and genotyping of ASF in real-time.

Virology↗

Partitioned Quantum Subspace Expansion

We present an iterative generalisation of the quantum subspace expansion algorithm used with a Krylov basis. The iterative construction connects a sequence of subspaces via their lowest energy states. Diagonalising a Hamiltonian in a given Krylov subspace requires the same quantum resources in both the single step and sequential cases. We propose a variance-based criterion for determining a good iterative sequence and provide numerical evidence that these good sequences display improved numerical stability over a single step in the presence of finite sampling noise. Implementing the generalisation requires additional classical processing with a polynomial overhead in the subspace dimension. By exchanging quantum circuit depth for additional measurements the quantum subspace expansion algorithm appears to be an approach suited to near term or early error-corrected quantum hardware. Our work suggests that the numerical instability limiting the accuracy of this approach can be substantially alleviated in a parameter-free way.

97 MATHEMATICS AND COMPUTING↗

NGPINT V3: a containerized orchestration Python software for discovery of next-generation protein–protein interactions

Abstract Summary Batch yeast two-hybrid (Y2H) assays, leveraged with next-generation sequencing, have afforded successful innovations for the analysis of protein–protein interactions. NGPINT is a Conda-based software designed to process the millions of raw sequencing reads resulting from Y2H–next-generation interaction screens. Over time, increasing compatibility and dependency issues have prevented clean NGPINT installation and operation. A system-wide update was essential to continue effective use with its companion software, Y2H-SCORES. We present NGPINT V3, a containerized implementation built with both Singularity and Docker, allowing accessibility across virtually any operating system and computing environment. Availability and implementation This update includes streamlined dependencies and container images hosted on Sylabs (https://cloud.sylabs.io/library/schuyler/ngpint/ngpint) and Dockerhub (https://hub.docker.com/r/schuylerds/ngpint), facilitating easier adoption and integration into high-throughput and cloud-computing workflows. Full instructions and software can be also found in the GitHub repository https://github.com/Wiselab2/NGPINT_V3 and Zenodo https://doi.org/10.5281/zenodo.15256036.

Biochemistry & Molecular Biology↗

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗