Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W↗

Packaging “vegetable oils”: Insights into plant lipid droplet proteins

Abstract Plant neutral lipids, also known as “vegetable oils”, are synthesized within the endoplasmic reticulum (ER) membrane and packaged into subcellular compartments called lipid droplets (LDs) for stable storage in the cytoplasm. The biogenesis, modulation, and degradation of cytoplasmic LDs in plant cells are orchestrated by a variety of proteins localized to the ER, LDs, and peroxisomes. Recent studies of these LD-related proteins have greatly advanced our understanding of LDs not only as steady oil depots in seeds but also as dynamic cell organelles involved in numerous physiological processes in different tissues and developmental stages of plants. In the past 2 decades, technology advances in proteomics, transcriptomics, genome sequencing, cellular imaging and protein structural modeling have markedly expanded the inventory of LD-related proteins, provided unprecedented structural and functional insights into the protein machinery modulating LDs in plant cells, and shed new light on the functions of LDs in nonseed plant tissues as well as in unicellular algae. Here, we review critical advances in revealing new LD proteins in various plant tissues, point out structural and mechanistic insights into key proteins in LD biogenesis and dynamic modulation, and discuss future perspectives on bridging our knowledge gaps in plant LD biology.

Cai, Yingqi (ORCID:0000000203575809)↗

Transduction-like gene transfer in the methanogen Methanococcus voltae

Strain PS of Methanococcus voltae (a methanogenic, anaerobic archaebacterium) was shown to generate spontaneously 4.4-kbp chromosomal DNA fragments that are fully protected from DNase and that, upon contact with a cell, transform it genetically. This activity, here called VTA (voltae transfer agent), affects all markers tested: three different auxotrophies (histidine, purine, and cobalamin) and resistance to BES (2-bromoethanesulfonate, an inhibitor of methanogenesis). VTA was most effectively prepared by culture filtration. This process disrupted a fraction of the M. voltae cells (which have only an S-layer covering their cytoplasmic membrane). VTA was rapidly inactivated upon storage. VTA particles were present in cultures at concentrations of approximately two per cell. Gene transfer activity varied from a minimum of 2 x 10(-5) (BES resistance) to a maximum of 10(-3) (histidine independence) per donor cell. Very little VTA was found free in culture supernatants. The phenomenon is functionally similar to generalized transduction, but there is no evidence, for the time being, of intrinsically viral (i.e., containing a complete viral genome) particles. Consideration of VTA DNA size makes the existence of such viral particles unlikely. If they exist, they must be relatively few in number;perhaps they differ from VTA particles in size and other properties and thus escaped detection. Digestion of VTA DNA with the AluI restriction enzyme suggests that it is a random sample of the bacterial DNA, except for a 0.9-kbp sequence which is amplified relative to the rest of the bacterial chromosome. A VTA-sized DNA fraction was demonstrated in a few other isolates of M. voltae.

Methanococcus/genetics/growth & development/metabo↗

Novosphingobium aromaticivorans LigR coordinates transcription of genes involved in metabolism of multiple types of aromatics

Aromatic compounds are a ubiquitous and diverse family of chemicals with functions as biomolecules, natural products, industrial chemicals, and pollutants. Novosphingobium aromaticivorans DSM 12444 uses multiple inducible pathways to catabolize H-, G-, and S-type aromatics that contain zero, one, or two methoxy groups, respectively. Here, we obtain a systems-level view of the transcriptional control of its aromatic metabolic pathways. Several in vitro analyses found that a N. aromaticivorans homolog of the Sphingobium lignivorans SYK-6 transcription factor LigR bound genomic DNA upstream of genes involved in metabolism of multiple aromatic types. We found that a ΔLigR mutant had growth defects on all three types of aromatics as sole carbon sources. Transcriptomic analysis revealed that LigR was required to increase expression of gene products that function in metabolism of all three aromatic types. We also found that, in media containing both glucose and an aromatic carbon source, the ΔLigR mutant directed intermediates through alternative aromatic metabolic pathways. Protein-DNA binding assays showed that N. aromaticivorans LigR binds immediately upstream of promoters of genes involved in aromatic metabolism. We found that N. aromaticivorans LigR coordinates the expression of enzymes that function in the catabolism of H-, G-, and S-type aromatics, and that there are differences in the role of LigR in N. aromaticivorans and S. lignivorans. A comparative genomic analysis predicted that LigR homologs and the aromatic-metabolizing genes that it directly regulates are often co-localized in the genomes of Sphingomonadales, but often not found in this arrangement in many other known aromatic metabolizing bacteria.

Aromatic Compound Degradation↗

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

Three pairs of fungal Trametes strains isolated from distinct geographic origins show conserved genomic features and adaptive response to plant biomass

The genomes of white-rot fungi hold extended repertoires of enzymes active on virtually all the chemical bonds that intertwine lignocellulose polymers, and several Trametes species have been identified as powerful tools for biorefinery or bioremediation. However, only few studies have addressed the intra-species polymorphism one would expect from fungal strains collected in contrasted environments. We compared the genome sequence of pairs of strains collected in different geographic areas, for each of three fungal species. Using an updated list of the predicted functions for fungal ligno- and cellulolytic enzymes (CAZymes), we observed a high conservation of the gene repertoires among the six strains. We compared the adaptative response of the fungi grown on crystalline cellulose, wheat straw, aspen or pine sawdust by transcriptomics and secretomics. The gene regulation profiles were determined by the species and the substrates, rather than the strain. The secretomes did not show marked differences in the sets of secreted CAZymes after 3 day-growth on the substrates. We identified five transcription factor genes and two sesquiterpenoid synthesis genes induced during growth on lignocellulose. Wider studies using larger sets of strains will be necessary to evaluate the genericity of our findings, and to assess the phenotype diversity one could expect from geographic diversity as compared to taxonomic diversity in Trametes fungi.

Drula, E. [French National Research Institute for ↗

Chance of Necessity: Modeling Origins of Life

The fundamental nature of processes that led to the emergence of life has been a subject of long-standing debate. One view holds that the origin of life is an event governed by chance, and the result of so many random events is unpredictable. This view was eloquently expressed by Jacques Monod in his book Chance or Necessity. In an alternative view, the origin of life is considered a deterministic event. Its details need not be deterministic in every respect, but the overall behavior is predictable. A corollary to the deterministic view is that the emergence of life must have been determined primarily by universal chemistry and biochemistry rather than by subtle details of environmental conditions. In my lecture I will explore two different paradigms for the emergence of life and discuss their implications for predictability and universality of life-forming processes. The dominant approach is that the origin of life was guided by information stored in nucleic acids (the RNA World hypothesis). In this view, selection of improved combinations of nucleic acids obtained through random mutations drove evolution of biological systems from their conception. An alternative hypothesis states that the formation of protocellular metabolism was driven by non-genomic processes. Even though these processes were highly stochastic the outcome was largely deterministic, strongly constrained by laws of chemistry. I will argue that self-replication of macromolecules was not required at the early stages of evolution; the reproduction of cellular functions alone was sufficient for self-maintenance of protocells. In fact, the precise transfer of information between successive generations of the earliest protocells was unnecessary and could have impeded the discovery of cellular metabolism. I will also show that such concepts as speciation and fitness to the environment, developed in the context of genomic evolution also hold in the absence of a genome.

Pohorille, Andrew↗

Colistin resistance plasmids dually enhance bacterial virulence and antibiotic resistance via surface polysaccharide biosynthesis

Plasmids carrying the mobilized colistin-resistance gene mcr-1 are prevalent among multidrug-resistant Gram-negative pathogens, yet their broad impact on bacterial physiology and virulence remains unclear. Here, we demonstrate that acquisition of an mcr-1 plasmid concurrently increases antimicrobial resistance and pathogenicity in Escherichia coli. On the same plasmid, the XRE-family transcriptional regulator EcaR cooperates with MCR-1 to activate the wec operon, driving biosynthesis of two surface polysaccharides: enterobacterial common antigen (ECA) and a high-molecular-weight O-chain. Expression of these surface polysaccharides increases bile resistance and virulence in a murine model and further elevates colistin resistance. MCR-1 enhances transcription of upstream genes in the wec operon, whereas EcaR directly activates an internal promoter (PwecE) to induce downstream gene expression. Thus, both components are required for surface polysaccharide expression, and deletion of either abolishes the phenotype. Genomic analysis of publicly available mcr plasmids reveals widespread co-occurrence of mcr-1 and ecaR on IncI2 and IncX4 plasmids, indicating their functional complementarity. These findings uncover a mechanism by which resistance plasmids remodel the bacterial surface, linking horizontal gene transfer to coordinated regulation of antimicrobial resistance and virulence.

Antimicrobial resistance↗

A constraint-based framework for exploring the impact of multireaction dependencies on metabolic functions

Abstract Metabolism operates under physico-chemical constraints that result in multireaction dependencies. Understanding how multireaction dependencies affect metabolic phenotypes remains challenging, hindering their biotechnological applications. Here, we propose the concept of a forcedly balanced complex that allows to efficiently determine the effects of specific multireaction dependencies on metabolic network functions in constrained-based models. Using this concept, we found that the fraction of multireaction dependencies induced by forcedly balanced complexes in genome-scale metabolic networks followed power law with exponential cut-off. We identified forcedly balanced complexes that are lethal in cancer but have little effect on growth in healthy tissue models. In addition, these forcedly balanced complexes are largely specific to models of particular cancer types. Therefore, multireaction dependencies resulting from forced balancing of complexes represent an innovative means to control cancers that, we argue, can be implemented via transporter engineering. The presented constraint-based approaches pave the way for using multireaction dependencies in metabolic engineering for diverse biotechnological applications.

Küken, Anika↗

Plant Metabolic Network 16: expansion of underrepresented plant groups and experimentally supported enzyme data

Abstract The Plant Metabolic Network (PMN) is a free online database of plant metabolism available at https://plantcyc.org. The latest release, PMN 16, provides metabolic databases representing >1200 metabolic pathways, 1.3 million enzymes, >8000 metabolites, >10 000 reactions and >15 000 citations for 155 plant and green algal genomes, as well as a pan-plant reference database called PlantCyc. This release contains 29 additional genomes compared with PMN 15, including species listed by the African Orphan Crop Consortium and nonflowering plant species. Furthermore, 52 new enzymes with experimentally supported function information have been included in this release. The single-species databases contain a combination of experimental information from the literature and computationally predicted information obtained through PMN’s database generation pipeline for a single species, while PlantCyc contains only experimental information but for any species within Viridiplantae. PMN is a comprehensive resource for querying, visualizing, analyzing and interpreting omics data with metabolic knowledge. It also serves as a useful and interactive tool for teaching plant metabolism.

Hawkins, Charles (ORCID:0000000312849047)↗

An overactivated ATR/CHK1 pathway is responsible for the prolonged G2 accumulation in irradiated AT cells

Induction of checkpoint responses in G1, S, and G2 phases of the cell cycle after exposure of cells to ionizing radiation (IR) is essential for maintaining genomic integrity. Ataxia telangiectasia mutated (ATM) plays a key role in initiating this response in all three phases of the cell cycle. However, cells lacking functional ATM exhibit a prolonged G2 arrest after IR, suggesting regulation by an ATM-independent checkpoint response. The mechanism for this ataxia telangiectasia (AT)-independent G2-checkpoint response remains unknown. We report here that the G2 checkpoint in irradiated human AT cells derives from an overactivation of the ATR/CHK1 pathway. Chk1 small interfering RNA abolishes the IR-induced prolonged G2 checkpoint and radiosensitizes AT cells to killing. These results link the activation of ATR/CHK1 with the prolonged G2 arrest in AT cells and show that activation of this G2 checkpoint contributes to the survival of AT cells.

NASA Discipline Radiation Health↗

Gene conversion is strongly induced in human cells by double-strand breaks and is modulated by the expression of BCL-x(L)

Homology-directed repair (HDR) of DNA double-strand breaks (DSBs) contributes to the maintenance of genomic stability in rodent cells, and it has been assumed that HDR is of similar importance in DSB repair in human cells. However, some outcomes of homologous recombination can be deleterious, suggesting that factors exist to regulate HDR. We demonstrated previously that overexpression of BCL-2 or BCL-x(L) enhanced the frequency of X-ray-induced TK1 mutations, including loss of heterozygosity events presumed to arise by mitotic recombination. The present study was designed to test whether HDR is a prominent DSB repair pathway in human cells and to determine whether ectopic expression of BCL-x(L) affects HDR. Using TK6-neo cells, we find that a single DSB in an integrated HDR reporter stimulates gene conversion 40-50-fold, demonstrating efficient DSB repair by gene conversion in human cells. Significantly, DSB-induced gene conversion events are 3-4-fold more frequent in TK6 cells that stably overexpress the antiapoptotic protein BCL-X(L). Thus, HDR plays an important role in maintaining genomic integrity in human cells, and ectopic expression of BCL-x(L) enhances HDR of DSBs. This is the first study to highlight a function for BCL-x(L) in modulating DSB repair in human cells.

NASA Discipline Radiation Health↗

Murine Host-gut Microbiota Interactions are Modulated During Spaceflight

The rodent habitat on the International Space Station has provided critical insight into the impact of spaceflight on mammalian physiology. These effects include dysfunction of carbohydrate, steroid and lipid metabolism, and immune response, as well as induction of symptoms characteristic of liver disease, insulin resistance, osteopenia and myopathy, which are anticipated to intensify over long-duration spaceflight. Although these physiological responses can involve the microbiome, the host-microorganism interactions during spaceflight are still largely unknown. NASA GeneLab curates a wide range of space research data and the current work harnesses GeneLab multi’omic data from recent Rodent Research studies to explore changes to gut microbiota during spaceflight and their associations with host physiology when compared to ground controls. Using a hybrid analysis of DNA barcoding and whole genome shotgun data, an array of bacteria, fungi and nematodes could be identified at species level, and significant differences in relative abundances associated with spaceflight. Functional prediction based on differential abundance of species and metagenome gene inventories as well as metatranscriptomic gene expression at the host-gut microbiome interface implicate microbiota interactions could contribute to spaceflight pathology. Harnessing carefully curated publicly available data, such as from Genelab, to generate multi‘omic space science discoveries can help decipher the complex host-microbiome interactions that influence both health on Earth and the feasibility of long-duration spaceflight.

Microbiome↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

An Open-Science Approach to Address Individual Response to Simulated GCR In Genetically Diverse Populations of Mice and Humans

This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm). A genome-wide association study (GWAS) was performed on the mouse strains between DNA damage responses to space radiation and single nucleotide polymorphisms (SNPs). We found SNPs, which were significantly associated to the RIF phenotype, mapped to genes and pathways that are functionally linked to health hazards for deep space exploration (e.g. carcinogenesis, nervous system damage and immune dysfunction). Some of these SNPs were located within protein coding regions, potentially interfering with protein functions and providing promising genetic targets for countermeasures. We also found correlations between both spontaneous and radiation-induced DNA damage and SNPs mapped to pathways associated with cellular metabolism. GWAS is undergoing for the human data. All data have been made available via the NASA Space Biology Open-Science database (genelab.nasa.gov) and we will discuss how various genomic and transcriptomic datasets can be accessed for modeling and integrated using machine learning methods for discovering new radiation biology.

Sylvain V Costes↗

Transcription of hepatitis B surface antigen shifts from cccDNA to integrated HBV DNA during treatment

The cornerstone of functional cure for chronic hepatitis B (CHB) is hepatitis B surface antigen (HBsAg) loss from blood. HBsAg is encoded by covalently closed circular DNA (cccDNA) and HBV DNA integrated into the host genome (iDNA). Nucleos(t)ide analogs (NUCs), the mainstay of CHB treatment, rarely lead to HBsAg loss, which we hypothesized was due to continued iDNA transcription despite decreased cccDNA transcription. To test this, we applied a multiplex droplet digital PCR that identifies the dominant source of HBsAg mRNAs to 3,436 single cells from paired liver biopsies obtained from 10 people with CHB and HIV receiving NUCs. With increased NUC duration, cells producing HBsAg mRNAs shifted their transcription from chiefly cccDNA to chiefly iDNA. This shift was due to both a reduction in the number of cccDNA-containing cells and diminished cccDNA-derived transcription per cell; furthermore, it correlated with reduced detection of proteins deriving from cccDNA but not iDNA. Despite this shift in the primary source of HBsAg, rare cells remained with detectable cccDNA-derived transcription, suggesting a source for maintaining the replication cycle. Functional cure must address both iDNA and residual cccDNA transcription. Further research is required to understand the significance of HBsAg when chiefly derived from iDNA.

59 BASIC BIOLOGICAL SCIENCES↗

A gene desert required for regulatory control of pleiotropic Shox2 expression and embryonic survival

Approximately a quarter of the human genome consists of gene deserts, large regions devoid of genes often located adjacent to developmental genes and thought to contribute to their regulation. However, defining the regulatory functions embedded within these deserts is challenging due to their large size. Here, we explore the cis-regulatory architecture of a gene desert flanking the Shox2 gene, which encodes a transcription factor indispensable for proximal limb, craniofacial, and cardiac pacemaker development. We identify the gene desert as a regulatory hub containing more than 15 distinct enhancers recapitulating anatomical subdomains of Shox2 expression. Ablation of the gene desert leads to embryonic lethality due to Shox2 depletion in the cardiac sinus venosus, caused in part by the loss of a specific distal enhancer. The gene desert is also required for stylopod morphogenesis, mediated via distributed proximal limb enhancers. In summary, our study establishes a multi-layered role of the Shox2 gene desert in orchestrating pleiotropic developmental expression through modular arrangement and coordinated dynamics of tissue-specific enhancers.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge

Understanding the interactions and regulatory relationships among biomolecules is essential for deciphering complex biological systems and elucidating the mechanisms behind diverse biological functions. Traditionally, the collection of such molecular interaction data has relied on expert curation, a process that is both time-consuming and labor-intensive. To address these limitations, this study explores the use of large language models (LLMs) to automate the genome-scale extraction of molecular interaction knowledge. Here, we evaluate the performance of various LLMs on key biological tasks, including the identification of protein-protein interactions, detection of genes associated with pathways influenced by low-dose radiation, and inference of gene regulatory relationships. Our findings demonstrate that larger LLMs tend to perform better, particularly in extracting intricate gene and protein interactions. Despite their strengths, these models face challenges in recognizing functionally diverse gene groups and highly correlated regulatory relationships. Through a comprehensive analysis using established molecular interaction and pathway databases, we show that LLMs possess the potential to identify relevant biomolecules and predict their interactions, offering valuable insights and marking a significant step toward AI-driven biological knowledge discovery.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗