Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Pioneer Venus spacecraft design and operation

The Pioneer Venus Orbiter and Multiprobe spacecraft design and operation enabled both remote and in-situ measurements of the Venusian environment from the outermost fringes of the atmosphere all the way to the surface. Both spacecraft were spin-stabilized and solar-cell powered from launch to Venus. Since orbit insertion, the Orbiter has been transmitting measurements from a highly elliptical 24-h orbit with periapsis altitudes down to about 150 km. Data rates up to 2048 bits/s have been utilized through a despun high-gain antenna transmitting at S-band frequency. Spacecraft attitudes, orbit periods, and periapsis altitudes are being maintained as required with a hydrazine propulsion system. The Multiprobe spacecraft (Bus with all four Probes attached) performed the necessary Probe checkouts and deployed the Probes to achieve the desired Probe and Bus targeting. Silver-zinc batteries provided the necessary power on each of the four Probes from separation from the Bus through the entry/descent sequence. Data rates of 256 and 128 bits/s on the Large Probe were maintained with 40-W radiated power, and 64 and 16 bits/s on the Small Probes were maintained with 10-W radiated power, through omni antennas directly to Earth-based stations. Each Probe's entry/descent sequence was controlled with a hardwired entry sequence programmer to achieve the desired scientific and spacecraft operations.

Nothwang, G. J.

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean

Using Fractional Clock-Period Delays in Telemetry Arraying

A set of special digital all-pass finite-impulse- response (FIR) filters produces phase shifts equivalent to delays that equal fractions of the sampling or clock period of a telemetry-data-processing system. These filters have been used to enhance the arraying of telemetry signals that have been received at multiple ground stations from spacecraft (see figure). Somewhat more specifically, these filters have been used to align, in the time domain, the telemetry-data sequences received by the various antennas, in order to maximize the signal-to-noise ratio of the composite telemetric signal obtained by summing the signals received by the antennas. The term arraying in this context denotes a method of enhanced reception of telemetry signals in which several antennas are used to track a single spacecraft. Each antenna receives a signal that comprises a sum of telemetry data plus noise, and these sum data are sent to an arraying combiner for processing. Correlation is the means used to align the set of data from one antenna with that from another antenna. After the data from all the antennas have been aligned in the time domain, they are all added together, sample by sample.

Fort, David

Multi-omics data resource: Data package 24 (Pck024)

The data package consists of isolated pancreatic islets from 3 human donors treated with IL-1β, IFNγ or IL-1β + IFNγ for 6 h and IL-1β, IFNγ, IL-1β + IFNγ, IL-1β + IFNγ + NMMA or NMMA for 18 h and submitted for scRNA-seq. This study examines cytokine-stimulated changes in gene expression in human islets using single-cell RNA sequencing. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE251730 Publication: 10.1093/function/zqae015

Sarkar, Soumyadeep [Pacific Northwest National Lab

Genomics Study of Effect of Redox-Active Metalloporphyrin on Murine Retina During Spaceflight

Astronauts returning from spaceflight have experienced eye problems, which may decrease retinal performance and lead to long-term effects on visual acuity. This study leverages the collected data from spaceflown murine retinas that were treated with redox-active metalloporphyrin (BuOE) to mitigate spaceflight-induced changes. 10-week-old adult C57BL/6 male mice (n=5 in each of BuOE treated and saline control groups) were flown on Space-X 24 to the ISS national lab, kept in low earth orbit for 35 days and returned to Earth alive. Our analysis of RNA-sequencing data generated from subsequent murine retina tissues uncovered genes, pathways, and epigenetic modifications consistent with therapeutic potential of BuOE. For spaceflown murine samples, the treatment group show differentially expressed genes relative to saline controls that reached significance (adjusted p-value < 0.05) and included genes Gpx3 and Crhbp, which are related to protection against cell oxidative damage and cellular response to organonitrogen compounds. Ranked fold-changes from the same contrast were used for gene set enrichment analysis, which showed biological processes reaching significance (adjusted p-value < 0.05) including glutathione metabolic processes and cellular response to xenobiotic stimulus. The findings from this investigation have the potential to provide valuable insights into the molecular mechanisms underlying conditions like spaceflight associated neuro-ocular syndrome and assess the effectiveness of BuOE as a countermeasure for astronauts experiencing neuro-ophthalmic abnormalities, which can lead to long-term effects on visual acuity.

Machine Learning

Genomics Study of Effect of Redox-Active Metalloporphyrin on Murine Retina During Spaceflight

Astronauts returning from spaceflight have experienced eye problems, which may decrease retinal performance and lead to long-term effects on visual acuity. This study leverages the collected data from spaceflown murine retinas that were treated with redox-active metalloporphyrin (BuOE) to mitigate spaceflight-induced changes. 10-week-old adult C57BL/6 male mice (n=5 in each of BuOE treated and saline control groups) were flown on Space-X 24 to the ISS national lab, kept in low earth orbit for 35 days and returned to Earth alive. Our analysis of RNA-sequencing data generated from subsequent murine retina tissues uncovered genes, pathways, and epigenetic modifications consistent with therapeutic potential of BuOE. For spaceflown murine samples, the treatment group show differentially expressed genes relative to saline controls that reached significance (adjusted p-value < 0.05) and included genes Gpx3 and Crhbp, which are related to protection against cell oxidative damage and cellular response to organonitrogen compounds. Ranked fold-changes from the same contrast were used for gene set enrichment analysis, which showed biological processes reaching significance (adjusted p-value < 0.05) including glutathione metabolic processes and cellular response to xenobiotic stimulus. The findings from this investigation have the potential to provide valuable insights into the molecular mechanisms underlying conditions like spaceflight associated neuro-ocular syndrome and assess the effectiveness of BuOE as a countermeasure for astronauts experiencing neuro-ophthalmic abnormalities, which can lead to long-term effects on visual acuity.

Machine Learning

GL4U: Training the next generation of bioinformaticians, one omics datatype at a time

Spaceflight modifies gene expression in every organism examined to date, including humans. Understanding how these gene expression changes affect physiology is crucial for the development of countermeasures to enable long-duration manned missions. NASA’s GeneLab project provides researchers open access to multi-omics data, including genetic and gene expression data, from spaceflight experiments that can be mined to understand the effects of spaceflight on biological systems. To ensure new knowledge generation through data re-use, it is important to maximize the number of scientists who utilize GeneLab data. Training students on the GeneLab platform is the best way to create long-term adopters of this NASA database and its tools. Turning students into future instructors and advocates will also accelerate the dissemination of these data and tools to the broader scientific community. Therefore, in collaboration with the GeneLab Educational Working Group (EWG), GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. During the bootcamp, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U direct training pilot program was conducted in June 2021 in collaboration with USRA and San Jose State University (SJSU). During the pilot, SJSU students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrates the capacity of GL4U for training young scientists and encouraging data re-use.

Jonathan Matthew Galazka

LinkFinder: An expert system that constructs phylogenic trees

An expert system has been developed using the C Language Integrated Production System (CLIPS) that automates the process of constructing DNA sequence based phylogenies (trees or lineages) that indicate evolutionary relationships. LinkFinder takes as input homologous DNA sequences from distinct individual organisms. It measures variations between the sequences, selects appropriate proportionality constants, and estimates the time that has passed since each pair of organisms diverged from a common ancestor. It then designs and outputs a phylogenic map summarizing these results. LinkFinder can find genetic relationships between different species, and between individuals of the same species, including humans. It was designed to take advantage of the vast amount of sequence data being produced by the Genome Project, and should be of value to evolution theorists who wish to utilize this data, but who have no formal training in molecular genetics. Evolutionary theory holds that distinct organisms carrying a common gene inherited that gene from a common ancestor. Homologous genes vary from individual to individual and species to species, and the amount of variation is now believed to be directly proportional to the time that has passed since divergence from a common ancestor. The proportionality constant must be determined experimentally; it varies considerably with the types of organisms and DNA molecules under study. Given an appropriate constant, and the variation between two DNA sequences, a simple linear equation gives the divergence time.

Inglehart, James

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection

Exploring Saccharomycotina Yeast Ecology Through an Ecological Ontology Framework

Yeasts in the subphylum Saccharomycotina are found across the globe in disparate ecosystems. A major aim of yeast research is to understand the diversity and evolution of ecological traits, such as carbon metabolic breadth, insect association, and cactophily. This includes studying aspects of ecological traits like genetic architecture or association with other phenotypic traits. Genomic resources in the Saccharomycotina have grown rapidly. Ecological data, however, are still limited for many species, especially those only known from species descriptions where usually only a limited number of strains are studied. Moreover, ecological information is recorded in natural language format limiting high throughput computational analysis. To address these limitations, we developed an ontological framework for the analysis of yeast ecology. A total of 1,088 yeast strains were added to the Ontology of Yeast Environments (OYE) and analyzed in a machine-learning framework to connect genotype to ecology. This framework is flexible and can be extended to additional isolates, species, or environmental sequencing data. Widespread adoption of OYE would greatly aid the study of macroecology in the Saccharomycotina subphylum.

59 BASIC BIOLOGICAL SCIENCES

Post-Flight Microbial Analysis of Samples from the International Space Station Water Recovery System and Oxygen Generation System

The Regenerative, Environmental Control and Life Support System (ECLSS) on the International Space Station (ISS) includes the the Water Recovery System (WRS) and the Oxygen Generation System (OGS). The WRS consists of a Urine Processor Assembly (UPA) and Water Processor Assembly (WPA). This report describes microbial characterization of wastewater and surface samples collected from the WRS and OGS subsystems, returned to KSC, JSC, and MSFC on consecutive shuttle flights (STS-129 and STS-130) in 2009-10. STS-129 returned two filters that contained fluid samples from the WPA Waste Tank Orbital Recovery Unit (ORU), one from the waste tank and the other from the ISS humidity condensate. Direct count by microscopic enumeration revealed 8.38 x 104 cells per mL in the humidity condensate sample, but none of those cells were recoverable on solid agar media. In contrast, 3.32 x lOs cells per mL were measured from a surface swab of the WRS waste tank, including viable bacteria and fungi recovered after S12 days of incubation on solid agar media. Based on rDNA sequencing and phenotypic characterization, a fungus recovered from the filter was determined to be Lecythophora mutabilis. The bacterial isolate was identified by rDNA sequence data to be Methylobacterium radiotolerans. Additional UPA subsystem samples were returned on STS-130 for analysis. Both liquid and solid samples were collected from the Russian urine container (EDV), Distillation Assembly (DA) and Recycle Filter Tank Assembly (RFTA) for post-flight analysis. The bacterium Pseudomonas aeruginosa and fungus Chaetomium brasiliense were isolated from the EDV samples. No viable bacteria or fungi were recovered from RFTA brine samples (N= 6), but multiple samples (N = 11) from the DA and RFTA were found to contain fungal and bacterial cells. Many recovered cells have been identified to genus by rDNA sequencing and carbon source utilization profiling (BiOLOG Gen III). The presence of viable bacteria and fungi from WRS and OGS subsystems demonstrates the need for continued monitoring of ECLSS during future ISS operations and investigation of advanced antimicrobial controls.

Birmele, Michele N.

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

NASA Program Fundamental Space Biology

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks

Ecological connectivity and habitat loss shape patterns of genetic diversity in a threatened salamander

Context The maintenance of genetic diversity is essential for preserving adaptive potential in populations, yet it is increasingly threatened by landscape alteration. The field of landscape genetics offers a framework for assessing how patch-level landscape conditions, modeled at multiple scales, influence genetic diversity. Objectives We sought to assess how local environmental features and connectivity influence genetic diversity across 74 four-toed salamander (Hemidactylium scutatum) breeding wetlands in the southeastern United States. Methods Using next-generation sequencing data and hierarchical Bayesian models, we examined genome-wide heterozygosity in relation to local landscape features and ecological connectivity. We also assessed the scale of effect of landscape features and tested for temporal lag effects. Results Genetic diversity was lower in wetlands with higher levels of historic deforestation and lower connectivity. An interaction between deforestation and connectivity indicated that deforestation had stronger negative effects in isolated wetlands but weaker effects in well-connected wetlands. Accounting for scale of effect and temporal lags was critical for detecting these relationships. Conclusions Our analyses highlight the importance of assessing the spatial scale (scale of effect) and temporal lag of landscape features to detect key drivers of genetic diversity. In line with population genetic theory, our results indicate that the genetic consequences of habitat loss do not affect populations uniformly and are most severe in isolated populations where gene flow cannot buffer against loss of diversity. Altogether, we highlight the importance of considering the interaction of habitat loss and connectivity in conservation genetic management.

Hemidactylium scutatum

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag

Identification and characterization of mono- and bifunctional galactan synthases in the pediatric pathogen Kingella kingae

The emerging pediatric pathogen Kingella kingae elaborates a lipopolysaccharide (LPS) that is extended with a galactofuranose homopolymer called galactan, which is a key virulence determinant that contributes to resistance to complement-mediated and neutrophil-mediated killing. Previous work has demonstrated that the pamABCDE locus is required for galactan synthesis. In this study, mutational studies suggested that the pamC gene product is a UDP-galactofuranose (Galf) transferase and is the galactan synthase. Analysis of genome sequence data revealed two distinct pamC alleles designated pamC1 and pamC2, which correlate with the two galactan structures in K. kingae. Examination of isogenic mutants expressing either pamC1 or pamC2 demonstrated that the pamC alleles are the determinants of galactan structure. Experiments with recombinant PamC1 and PamC2 in vitro established that these proteins are galactan synthases capable of extending synthetic Galf disaccharide acceptors in the presence of UDP-Galf. Homology analysis identified critical amino acids that are essential for PamC1 and PamC2 enzymatic activity both in vitro and in K. kingae. Structural analysis of the in vitro-modified synthetic acceptors implicated PamC1 as a monofunctional enzyme capable of generating a β-(1 → 5) Galf linkage and PamC2 as a bifunctional enzyme capable of generating β-(1 → 3) and β-(1 → 6) Galf linkages. This study advances our understanding of the GT2 family of UDP-galactofuranosyltransferases.

60 APPLIED LIFE SCIENCES

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING