Search NASA⌕ Search

SEARCH · Search NASA

Results for “annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Acute wood smoke exposure is associated with cell-specific hippocampal transcriptomic responses in an accelerated ovarian failure mouse model

Background Wildfire events are increasing in frequency and intensity, and aging individuals demonstrate heightened biological susceptibility to air pollution exposures including increased risk of neurological sequelae. Declining ovarian hormones levels that occur with aging in females along with associated systemic physiological and inflammatory changes may contribute to increased cerebral vulnerability to air pollution, representing a potential but underexplored mechanism. Menopause and the menopausal transition represent a period of profound physiological change that affects cardiovascular, neurological, and immune health. Methods We tested whether peri-menopausal–like hormonal status amplifies hippocampal responses to acute wood smoke (WS) using an ovary-intact, 4-vinylcyclohexene diepoxide (VCD) model of moderate accelerated ovarian failure (AOF) in female C57BL/6 mice. Animals were exposed to HEPA-filtered air (FA) or WS for 4 h/day over 2 consecutive days (∼0.5 mg/m³). Exposure characterization confirmed a complex mixture of combustion products with significant levels of both trace metals and gas release during WS exposure. Results Spatial transcriptomics (10x Visium; n = 4 sections/group) with automated cell-type annotation identified astrocytes, GABAergic and glutamatergic neurons, oligodendrocytes, revealed cell type-specific transcriptional alterations following WS exposure. Distinct transcriptional patterns were observed across all identified neuronal and glial cell populations. Conclusion Together, these findings define a cell-type specific transcriptomic framework describing how WS exposure and ovarian hormone decline interact to influence hippocampal responses and identify potential cellular pathways relevant to hippocampal vulnerability.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

Phosphate amendment drives bloom of RNA viruses after soil wet-up

Soil rewetting after a dry period results in a surge of activity and succession in both microbial and DNA virus communities. Less is known about the response of RNA viruses to soil rewetting—while they are highly diverse and widely distributed in soil, they remain understudied. We hypothesized that RNA viruses would show temporal succession following rewetting and that phosphate amendment would influence their trajectory, as viral proliferation may cause phosphorus limitation. Using 39 time-resolved metatranscriptomes and amplicon data, 2190 RNA viral populations were identified across five phyla, with 26 % of these predicted to infect bacteria, and 11 % fungi. Only 1.2 % of viral populations had annotated capsid genes, suggesting most persist via intracellular replication without a free virion phase. Phosphate amendment altered RNA viral community composition within the first week and amended vs. unamended communities remained distinguishable for up to three weeks. While the overall host community remained stable, certain bacterial populations showed reduced abundance in phosphate-amended soils, likely due to increased viral lysis, as RNA bacteriophages proliferated significantly. Notably, 60 % of the viruses with increased abundance under phosphate amendment belonged to basal Lenarviricota clades rather than well-known groups like Leviviricetes. We estimate RNA bacteriophage infections may affect 10 7 –10 9 bacteria per gram of soil, aligning with the total bacterial population (10 7 –10 10 g -1 soil), suggesting that RNA phages significantly influence bacterial communities post-wet-up, with phosphorus availability modulating this effect.

59 BASIC BIOLOGICAL SCIENCES↗

Root exudate lipids: Uncovering chemodiversity and carbon stability potential

Root-derived carbon has been shown to contribute more to soil carbon stocks than aboveground litter. Yet the molecular chemodiversity of root exudates remains poorly understood due to limited characterization and annotation. In this study, we characterized the molecular chemodiversity and production of metabolites and lipids in root exudates from field grown mature tall wheatgrass (Thinopyrum ponticum). We discovered a diversity of lipids, including substantial levels of triacylglycerols (∼19 μg/g fresh root per min), fatty acyls, sphingolipids, sterol lipids, and glycerophospholipids, some of which have not been previously documented in root exudates. By integrating tandem mass spectral library searching and deep learning-based chemical class assignment, our metabo-lipidomics approach significantly expanded the known molecular diversity of root exudates. Rates of lipid derived carbon production were approximately double that of polar metabolites (lipids: 81.52 ± 13.81 vs polar metabolites: 38.41 ± 5.93 μg C g −1 fresh root mass min −1 ) with an order of magnitude higher carbon to nitrogen ratios (lipids: 459 ± 90 vs polar metabolites: 14.40 ± 0.58). Exudate lipids displayed highly negative nominal oxidation state of carbon (−1.182 to −1.909), indicating that these compounds may be less favorable for microbial decomposition. Together our results suggest the potential of root exudate lipids to contribute to stable carbon pools in soil, supporting long-term carbon storage. This work advances understanding of plant-derived lipid inputs to soil and underscores the need for future studies on the functional roles of lipids in shaping root-microbe-soil interactions, microbial activity, soil structure, and nutrient availability – contributing to soil health.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Morphotype-resolved characterization of microalgal communities in a nutrient recovery process with ARTiMiS flow imaging microscopy

Microalgae-driven nutrient recovery represents a promising technology for phosphorus removal from wastewater while simultaneously generating biomass that can be valorized to offset treatment costs. As full-scale processes come online, system parameters including biomass composition must be carefully monitored to optimize performance and prevent culture crashes. In this study, flow imaging microscopy (FIM) was leveraged to characterize microalgal community composition in near real-time at a full-scale municipal wastewater treatment plant (WWTP) in Wisconsin, USA, and population and morphotype dynamics were examined to identify relationships between water chemistry, biomass composition, and system performance. Two FIM technologies, FlowCam and ARTiMiS, were evaluated as monitoring tools. ARTiMiS provided a more accurate estimate of total system biomass, and estimates derived from particle area as a proxy for biovolume yielded better approximations than particle counts. Deep learning classification models trained on annotated image libraries demonstrated equivalent performance between FlowCam and ARTiMiS, and convolutional neural network (CNN) classifiers proved significantly more accurate when compared to feature table-based dense neural network (DNN) models. Across a two-year study period, Scenedesmus spp. appeared most important for phosphorus removal, and were negatively impacted by elevated temperatures and increase in nitrite/nitrate concentrations. Chlorella and Monoraphidium also played an important role in phosphorus removal. For both Scenedesmus and Chlorella, smaller morphological types were more often associated with better system performance, whereas larger morphotypes likely associated with stress response(s) correlated with poor phosphorus recovery rates. Furthermore, these results demonstrate the potential of FIM as a critical technology for high-resolution characterization of industrial microalgal processes.

59 BASIC BIOLOGICAL SCIENCES↗

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics↗

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

High-Resolution Tandem Mass Spectrometry-Based Analysis of Model Lignin–Iron Complexes: Novel Pipeline and Complex Structures

Understanding the chemical nature of soil organic carbon (SOC) with great potential to bind iron (Fe) minerals is critical for predicting the stability of SOC. Organic ligands of Fe are among the top candidates for SOCs able to strongly sorb on Fe minerals, but most of them are still molecularly uncharacterized. To shed insights into the chemical nature of organic ligands in soil and their fate, this study developed a protocol for identifying organic ligands using ultrahigh-performance liquid chromatography-high-resolution tandem mass spectrometry (UHPLC-HRMS/MS) and metabolomic tools. The protocol was used for investigating the Fe complexes formed by model compounds of lignin-derived organic ligands, namely, caffeic acid (CA), p-coumaric acid (CMA), vanillin (VNL), and cinnamic acid (CNA). Isotopologue analysis of 54/56 Fe was used to screen out the potential UHPLC-HRMS (m/z) features for complexes formed between organic ligands and Fe, with multiple features captured for CA, CMA, VNL, and CNA when 35/37 Cl isotopologue analysis was used as supplementary evidence for the complexes with Cl. MS/MS spectra, fragment analysis, and structure prediction with SIRIUS were used to annotate the structures of mono/bidentate mono/biligand complexes. The analysis determined the structures of monodentate and bidentate complexes of FeL x Cl y (L: organic ligand, x = 1–4, y = 0–3) formed by model compounds. The protocol developed in this study can be used to identify unknown organic ligands occurring in complex environmental samples and shed light on the molecular-level processes governing the stability of the SOC.

54 ENVIRONMENTAL SCIENCES↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA↗

Evaluation of a Reference-Free Collision Cross Section Calibration Strategy for Proteomics Using SLIM-Based High-Resolution Ion Mobility Spectrometry–Mass Spectrometry

Ion mobility spectrometry (IMS) is a gas-phase analytical technique that separates ions with different sizes and shapes and is compatible with mass spectrometry (MS) to provide an additional separation dimension. The rapid nature of the IMS separation combined with the high sensitivity of MS-based detection and the ability to derive structural information on analytes in the form of the property collision cross section (CCS) makes IMS particularly well-suited for characterizing complex samples in -omics applications. In such applications, the quality of CCS from IMS measurements is critical to confident annotation of the detected components in the complex -omics samples. However, most IMS instrumentation in mainstream use requires calibration to calculate CCS from measured arrival times, with the most notable exception being drift tube IMS measurements using multifield methods. The strategy for calibrating CCS values, particularly selection of appropriate calibrants, has important implications for CCS accuracy, reproducibility, and transferability between laboratories. The conventional approach to CCS calibration involves explicitly defining calibrants ahead of data acquisition and crucially relies upon availability of reference CCS values. In this work, we present a novel reference-free approach to CCS calibration which leverages trends among putatively identified features and computational CCS prediction to conduct calibrations post-data acquisition and without relying on explicitly defined calibrants. We demonstrated the utility of this reference-free CCS calibration strategy for proteomics application using high-resolution structures for lossless ion manipulations (SLIM)-based IMS-MS. In conclusion, we first validated the accuracy of CCS values using a set of synthetic peptides and then demonstrated using a complex peptide sample from cell lysate.

59 BASIC BIOLOGICAL SCIENCES↗

Central transcriptional regulator controls photosynthetic growth and carbon storage in response to high light

Carbon capture and biochemical storage are some of the primary drivers of photosynthetic yield and productivity. To elucidate the mechanisms governing carbon allocation, we designed a photosynthetic light response test system for genetic and metabolic carbon assimilation tracking, using microalgae as simplified plant models. The systems biology mapping of high light-responsive photophysiology and carbon utilization dynamics between two variants of the same Picochlorum celeri species, TG1 and TG2 elucidated metabolic bottlenecks and transport rates of intermediates using instationary 13 C-fluxomics. Simultaneous global gene expression dynamics showed 73% of the annotated genes responding within one hour, elucidating a singular, diel-responsive transcription factor, closely related to the CCA1/LHY clock genes in plants, with significantly altered expression in TG2. Transgenic P. celeri TG1 cells expressing the TG2 CCA1/LHY gene, showed 15% increase in growth rates and 25% increase in storage carbohydrate content, supporting a coordinating regulatory function for a single transcription factor.

09 BIOMASS FUELS↗

Barcoded overexpression screens in gut Bacteroidales identify genes with roles in carbon utilization and stress resistance

Abstract A mechanistic understanding of host-microbe interactions in the gut microbiome is hindered by poorly annotated bacterial genomes. While functional genomics can generate large gene-to-phenotype datasets to accelerate functional discovery, their applications to study gut anaerobes have been limited. For instance, most gain-of-function screens of gut-derived genes have been performed in Escherichia coli and assayed in a small number of conditions. To address these challenges, we develop Barcoded Overexpression BActerial shotgun library sequencing (Boba-seq). We demonstrate the power of this approach by assaying genes from diverse gut Bacteroidales overexpressed in Bacteroides thetaiotaomicron . From hundreds of experiments, we identify new functions and phenotypes for 29 genes important for carbohydrate metabolism or tolerance to antibiotics or bile salts. Highlights include the discovery of a d -glucosamine kinase, a raffinose transporter, and several routes that increase tolerance to ceftriaxone and bile salts through lipid biosynthesis. This approach can be readily applied to develop screens in other strains and additional phenotypic assays.

59 BASIC BIOLOGICAL SCIENCES↗

Structural diversity and clustering of bacterial flagellar outer domains

Supercoiled flagellar filaments function as mechanical propellers within the bacterial flagellum complex, playing a crucial role in motility. Flagellin, the building block of the filament, features a conserved inner D0/D1 core domain across different bacterial species. In contrast, approximately half of the flagellins possess additional, highly divergent outer domain(s), suggesting varied functional potential. In this study, we report atomic structures of flagellar filaments from three distinct bacterial species: Cupriavidus gilardii , Stenotrophomonas maltophilia , and Geovibrio thiophilus . Our findings reveal that the flagella from the facultative anaerobic G. thiophilus possesses a significantly more negatively charged surface, potentially enabling adhesion to positively charged minerals. Furthermore, we analyze all AlphaFold predicted structures for annotated bacterial flagellins, categorizing the flagellin outer domains into 682 structural clusters. This classification provides insights into the prevalence and experimental verification of these outer domains. Remarkably, two of the flagellar structures reported herein belong to a distinct cluster, indicating additional opportunities on the study of the functional diversity of flagellar outer domains. Our findings underscore the complexity of bacterial flagellins and open up possibilities for future studies into their varied roles beyond motility.

Science & Technology - Other Topics↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics↗