Search NASA⌕ Search

SEARCH · Search NASA

Results for “Genetics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Phylogenomics and genetic analysis of solvent-producing Clostridium species

Abstract The genus Clostridium is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 C. beijerinckii , 57 C. saccharobutylicum , 4 C. saccharoperbutylacetonicum , 5 C. butyricum , 7 C. acetobutylicum , and 3 C. tetanomorphum genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or Escherichia coli for enhanced chemical production. Recently, gene sequences from this collection were used to engineer Clostridium autoethanogenum , a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production, as well as butanol, butanoic acid, hexanol and hexanoic acid production.

59 BASIC BIOLOGICAL SCIENCES↗

Applying MALDI-TOF MS to resolve morphologic and genetic similarities between two Dermacentor tick species of public health importance

Abstract Hard ticks (Acari: Ixodidae) have been historically identified by morphological methods which require highly specialized expertise and more recently by DNA-based molecular assays that involve high costs. Although both approaches provide complementary data for tick identification, each method has limitations which restrict their use on large-scale settings such as regional or national tick surveillance programs. To overcome those obstacles, the matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) has been introduced as a cost-efficient method for the identification of various organisms, as it balances performance, speed, and high data output. Here we describe the use of this technology to validate the distinction of two closely relatedDermacentortick species based on the development of the first nationwide MALDI-TOF MS reference database described to date. The dataset obtained from this protein-based approach confirms that tick specimens collected from United States regions west of the Rocky Mountains and identified previously asDermacentor variabilisare the recently described species,Dermacentor similis. Therefore, we propose that this integrative taxonomic tool can facilitate vector and vector-borne pathogen surveillance programs in the United States and elsewhere.

Science & Technology - Other Topics↗

Genetic control of morphological transitions in a coacervating protein template

Nature routinely exploits liquid–liquid phase separation (LLPS) of proteins to control the assembly and mineralization of hybrid materials. Here, we show that fusion of the Car9 silica-binding peptide to an elastin-like polypeptide (ELP) yields temperature- and sequence-programmable soft matter templates for the synthesis of silicified architectures ranging in size from nanometers to micrometers. Specifically, we demonstrate unprecedented control over the diameter of silica nanoparticles (SiNP) in the 30–60 nm range with 4 nm precision, show that a single arginine residue (R4) in the Car9 sequence underpins the transition from micelles to proteinosomes, and find that substitutions in other basic residues modulate electrostatic repulsion and solvation to enable access to kinetically trapped species. These structures, which include interconnected micelles, small (∼200 nm) and large (>5 µm) vesicles, are readily visualized by SEM imaging following silicification. Molecular dynamics (MD) simulations and AlphaFold predictions reveal that mutations in positively charged residues alter interfacial packing, hydration, and conformational freedom of the silica-binding segments. Overall, our results establish sequence and thermal energy as synergistic levers for morphological control across length scales using solid-binding ELPs and establish mineralization as a powerful tool to visualize the structure of dynamic soft matter assemblies.

hierarchy↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Genetically manipulated chloroplast stromal phosphate levels alter photosynthetic efficiency

Abstract The concentration of inorganic phosphate (Pi) in the chloroplast stroma must be maintained within narrow limits to sustain photosynthesis and to direct the partitioning of fixed carbon. However, it is unknown if these limits or the underlying contributions of different chloroplastic Pi transporters vary throughout the photoperiod or between chloroplasts in different leaf tissues. To address these questions, we applied live Pi imaging to Arabidopsis (Arabidopsis thaliana) wild-type plants and 2 loss-of-function transporter mutants: triose phosphate/phosphate translocator (tpt), phosphate transporter 2;1 (pht2;1), and tpt pht2;1. Our analyses revealed that stromal Pi varies spatially and temporally, and that TPT and PHT2;1 contribute to Pi import with overlapping tissue specificities. Further, the series of progressively diminished steady-state stromal Pi levels in these mutants provided the means to examine the effects of Pi on photosynthetic efficiency without imposing nutritional deprivation. ΦPSII and nonphotochemical quenching (NPQ) correlated with stromal Pi levels. However, the proton efflux activity of the ATP synthase (gH+) and the thylakoid proton motive force (pmf) were unaltered under growth conditions, but were suppressed transiently after a dark to light transition with return to wild-type levels within 2 min. These results argue against a simple substrate-level limitation of ATP synthase by depletion of stromal Pi, favoring more integrated regulatory models, which include rapid acclimation of thylakoid ATP synthase activity to reduced Pi levels.

54 ENVIRONMENTAL SCIENCES↗

High-throughput genetics enables identification of nutrient utilization and accessory energy metabolism genes in a model methanogen

Archaea are widespread in the environment and play fundamental roles in diverse ecosystems; however, characterization of their unique biology requires advanced tools. This is particularly challenging when characterizing gene function. Here, we generate randomly barcoded transposon libraries in the model methanogenic archaeon Methanococcus maripaludis and use high-throughput growth methods to conduct fitness assays (RB-TnSeq) across over 100 unique growth conditions. Using our approach, we identified new genes involved in nutrient utilization and response to oxidative stress. We identified novel genes for the usage of diverse nitrogen sources in M. maripaludis including a putative regulator of alanine deamination and molybdate transporters important for nitrogen fixation. Furthermore, leveraging the fitness data, we inferred that M. maripaludis can utilize additional nitrogen sources including $\tiny{L}$-glutamine, $\tiny{D}$-glucuronamide, and adenosine. Under autotrophic growth conditions, we identified a gene encoding a domain of unknown function (DUF166) that is important for fitness and hypothesize that it has an accessory role in carbon dioxide assimilation. Finally, comparing fitness costs of oxygen versus sulfite stress, we identified a previously uncharacterized class of dissimilatory sulfite reductase-like proteins (Dsr-LP; group IIId) that is important during growth in the presence of sulfite. When overexpressed, Dsr-LP conferred sulfite resistance and enabled use of sulfite as the sole sulfur source. The high-throughput approach employed here allowed for generation of a large-scale data set that can be used as a resource to further understand gene function and metabolism in the archaeal domain.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic and biochemical characterization of a radical SAM enzyme required for post-translational glutamine methylation of methyl-coenzyme M reductase

ABSTRACT Methyl-coenzyme M reductase (MCR), the key catalyst in the anoxic production and consumption of methane, contains an unusual 2-methylglutamine residue within its active site. In vitro data show that a B12-dependent radical SAM (rSAM) enzyme, designated MgmA, is responsible for this post-translational modification (PTM). Here, we show that two different MgmA homologs are able to methylate MCR in vivo when expressed in Methanosarcina acetivorans , an organism that does not normally possess this PTM. M. acetivorans strains expressing MgmA showed small, but significant, reductions in growth rates and yields on methylotrophic substrates. Structural characterization of the Ni(II) form of Gln-methylated M. acetivorans MCR revealed no significant differences in the protein fold between the modified and unmodified enzyme; however, the purified enzyme contained the heterodisulfide reaction product, as opposed to the free cofactors found in eight prior M. acetivorans MCR structures, suggesting that substrate/product binding is altered in the modified enzyme. Structural characterization of MgmA revealed a fold similar to other B12-dependent rSAMs, with a wide active site cleft capable of binding an McrA peptide in an extended, linear conformation. IMPORTANCE Methane plays a key role in the global carbon cycle and is an important driver of climate change. Because MCR is responsible for nearly all biological methane production and most anoxic methane consumption, it plays a major role in setting the atmospheric levels of this important greenhouse gas. Thus, a detailed understanding of this enzyme is critical for the development of methane mitigation strategies.

Rodriguez Carrero, Roy J. (ORCID:0000000184475641)↗

Genetic and metabolic drivers of membrane remodeling in Clostridium thermocellum under alcohol stress

Clostridium thermocellum is a leading candidate for consolidated bioprocessing of lignocellulosic biomass into biofuels due to its native cellulolytic capabilities. Beyond ethanol, C. thermocellum is being developed as a platform for producing higher-chain alcohols such as isobutanol and n-butanol. However, its physiological adaptations to alcohol stress remain poorly understood. Here, we investigate how C. thermocellum remodels its membrane lipid composition in response to exogenous ethanol, n-butanol, isobutanol, and butyrate. Exposure to linear alcohols such as n-butanol or to organic acids like butyrate increased the proportion of straight-chain fatty acids in the membrane at the expense of branched-chain species, whereas exposure to the branched alcohol isobutanol produced the opposite effect. Isotope tracer experiments demonstrated that C. thermocellum directly incorporates the carbon backbones of exogenous alcohols and acids into fatty acids, providing a mechanistic basis for these contrasting shifts. We show that the bifunctional aldehyde/alcohol dehydrogenase AdhE is essential for the assimilation of exogenous alcohols into fatty acids, acting through its oxidative activity by first oxidizing alcohols to aldehydes and then converting them to acyl-CoA intermediates. Deletion of the pyruvate:ferredoxin oxidoreductase isozyme pfor4 abolished branched-chain fatty acid synthesis, but supplementation with isobutanol restored production, indicating that Pfor4 substitutes for the canonical branched-chain α-keto acid dehydrogenase complex. These findings reveal two distinct routes for branched-chain fatty acid production in C. thermocellum: a Pfor4-dependent pathway from α-keto acid intermediates derived from amino acid synthesis, and an AdhE-dependent salvage pathway that assimilates exogenous branched-chain alcohols.

Acetivibrio thermocellus↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗