Search NASA⌕ Search

SEARCH · Search NASA

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Enhanced Spatial Proteomics and Metabolomics from a Single Tissue Section Using MALDI-MSI and LCM-microPOTS Platforms

Spatially resolved mass spectrometry (MS)-based multi-omics workflows are becoming more utilized for revealing the complex biology that occurs within tissues. However, these approaches commonly require multiple independent tissue sections to analyze the metabolite and protein compositions of these samples. This poses a significant challenge in preserving cell- or region-specific molecular fidelity, as variations between tissue sections can compromise the accurate correlation of molecular data. Here, in this study, we developed workflows for comprehensive multi-omics profiling from a single tissue section (STS) using different MS modalities. We enhanced the functionality of an electrically insulated substrate by employing metal-assisted approaches that enabled both MS-based untargeted spatial metabolomics and proteomics from STS. This allowed metabolite imaging using matrix-assisted laser desorption/ionization-MS imaging (MALDI-MSI), without compromising it for subsequent proteome profiling with laser capture microdissection (LCM)-based technology. Specifically, implementing copper tape as a backing for polyethylene naphthalate (PEN) slides enabled the detection of >140 metabolites across a poplar root tissue section using MALDI-trapped ion mobility spectrometry time of flight (timsTOF)-MS. Afterwards, we detected 6,571 unique proteins from two distinct root regions by leveraging LCM technology coupled to our microdroplet based sample preparation approach. We also developed an alternative workflow utilizing gold-coated PEN substrates for imaging with MALDI-Fourier-transform ion cyclotron resonance (FTICR)-MS, which permitted the profiling of >170 metabolites and the identification of 6,542 unique proteins across a single poplar root tissue section. These results were comparable to using each assay independently without modifications. These approaches offer new opportunities for high-resolution molecular profiling of multiple omics-levels across biological tissues.

Veličković, Marija [Pacific Northwest National Lab↗

Omics-driven onboarding of the carotenoid producing red yeast Xanthophyllomyces dendrorhous CBS 6938

Transcriptomics is a powerful approach for functional genomics and systems biology, yet it can also be used for genetic part discovery. Here, we derive constitutive and light-regulated promoters directly from transcriptomics data of the basidiomycete red yeast Xanthophyllomyces dendrorhous CBS 6938 (anamorph Phaffia rhodozyma) and use these promoters with other genetic elements to create a modular synthetic biology parts collection for this organism. X. dendrorhous is currently the sole biotechnologically relevant yeast in the Tremellomycete class-it produces large amounts of astaxanthin, especially under oxidative stress and exposure to light. Thus, we performed transcriptomics on X. dendrorhous under different wavelengths of light (red, green, blue, and ultraviolet) and oxidative stress. Differential gene expression analysis (DGE) revealed that terpenoid biosynthesis was primarily upregulated by light through crtI, while oxidative stress upregulated several genes in the pathway. Further gene ontology (GO) analysis revealed a complex survival response to ultraviolet (UV) where X. dendrorhous upregulates aromatic amino acid and tetraterpenoid biosynthesis and downregulates central carbon metabolism and respiration. The DGE data was also used to identify 26 constitutive and regulated genes, and then, putative promoters for each of the 26 genes were derived from the genome. Simultaneously, a modular cloning system for X. dendrorhous was developed, including integration sites, terminators, selection markers, and reporters. Each of the 26 putative promoters were integrated into the genome and characterized by luciferase assay in the dark and under UV light. The putative constitutive promoters were constitutive in the synthetic genetic context, but so were many of the putative regulated promoters. Notably, one putative promoter, derived from a hypothetical gene, showed ninefold activation upon UV exposure. Thus, this study reveals metabolic pathway regulation and develops a genetic parts collection for X. dendrorhous from transcriptomic data. Therefore, this study demonstrates that combining systems biology and synthetic biology into an omics-to-parts workflow can simultaneously provide useful biological insight and genetic tools for nonconventional microbes, particularly those without a related model organism. This approach can enhance current efforts to engineer diverse microbes.

60 APPLIED LIFE SCIENCES↗

Endurance exercise elicits temporal and sexual dimorphic multi-omics remodeling of liver metabolism revealed by MoTrPAC

The mechanisms by which exercise modulate liver metabolism, a central regulator of systemic metabolism, are poorly understood. Leveraging data from MoTrPAC, we analyzed liver adaptations across 1, 2, 4, and 8 weeks of exercise in male and female rats using multi-omic approaches. Female livers displayed a progressive increase in oxidative phosphorylation (OXPHOS) complexes (at the protein level), while male livers showed an increase in acetylation of OXPHOS, TCA cycle, and fatty acid oxidation enzymes. Exercise also enhanced liver cholesterol and bile acid synthesis, reducing liver lipid metabolites in males after 8 weeks of exercise. Male rats had higher fecal cholesterol and cholic acid levels, indicating a sex-specific mechanism of lipid excretion with exercise. Moreover, 8 weeks of training reduced markers related to hepatic stellate cell activation and fibrosis in both sexes. This study highlights the sexual dimorphic and temporal molecular signatures by which exercise modulates liver metabolism to provide hepatoprotective effects.

Kelty, Taylor↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

Novel CHI3L1 ‐Associated Angiogenic Phenotypes Define Glioma Microenvironments: Insights From Multi‐Omics Integration

ABSTRACT The CHI3L1 signaling pathway significantly influences glioma angiogenesis, but its role in the tumor microenvironment (TME) remains elusive. We propose a novelCHI3L1‐associated vascular phenotype classification for glioma through integrative analyses of multiple datasets with bulk and single‐cell transcriptome, genomics, digital pathology, and clinical data. We investigated the biological characteristics, genomic alterations, therapeutic vulnerabilities, and immune profiles within these phenotypes through a comprehensive multi‐omics approach. We constructed the vascular‐related risk (VR) score based onCHI3L1‐associated vascular signatures (CAVS) identified by machine learning algorithms. Utilizing unsupervised consensus clustering, gliomas were stratified into three distinct vascular phenotypes: Cluster A, marked by high vascularization and stromal activation with a relatively low levels of tumor‐infiltrating lymphocytes (TILs); Cluster B, characterized by moderate vascularization and stromal activity, coupled with a high density of TILs; and Cluster C, defined by low vascularization and sparse immune cell infiltration. We observed that the CAVS effectively indicated glioma‐associated angiogenesis and immune suppression by single‐cell RNA‐seq analysis. Moreover, the high‐VR‐score group exhibited enhanced angiogenic activity, reduced immune response, resistance to immunotherapy, and poorer clinical outcomes. The VR score independently predicted glioma prognosis and, combined with a nomogram, provided a robust clinical decision‐making tool. Potential drug prediction based on transcription factors for high‐risk patients was also performed. Our study reveals thatCHI3L1‐associated vascular phenotypes shape distinct immune landscapes in gliomas, offering insights for optimizing therapeutic strategies to improve patient outcomes.

Oncology↗

Multi-omics of a model bacterial consortium deciphers details of chitin decomposition in soil

Soil microorganisms interact to carry out decomposition of complex organic carbon and nitrogen compounds, such as chitin, but the high diversity and complexity of the soil microbiome and habitat have posed a challenge to elucidating such interactions. Here, we sought to address this challenge by analysis of a model soil consortium (MSC-2) consisting of eight soil bacterial species. Our aim was to elucidate the specific roles of the member species during chitin metabolism. Samples were collected from MSC-2 incubated in chitin-enriched soil over 3 months. Multi-omics was used to understand how the community composition, transcripts, proteins, and chitin decomposition shifted over time. The data clearly and consistently revealed a temporal shift during chitin decomposition with defined contributions by individual species. A Streptomyces genus member (sp001905665) was a key player in early steps of chitin decomposition, with other MSC-2 members being central in carrying out later steps. These results illustrate how multi-omics applied to a defined consortium untangles the interactions between soil microorganisms.

chitin↗

EVT 16s Data and Large Supplementary Files

Soil microorganisms often interact to carry out decomposition of complex organic carbon and nitrogen compounds, such as chitin, but the high diversity and complexity of the soil microbiome and habitat has posed a challenge to elucidating such interactions between soil microorganisms. Here, we seek to address this challenge through analysis of a model soil consortium (MSC-2) of eight soil bacterial species. Our aim was to elucidate specific roles of the member species during chitin metabolism. Samples were collected from MSC-2 incubated in chitin-enriched soil over three months. Multi-omics was used to understand how the community composition, transcripts, proteins and chitin decomposition shifted over time. The data clearly and consistently revealed a temporal shift during chitin decomposition with defined contributions by individual species. A Streptomyces genus member (sp001905665) was a key player in early steps of chitin decomposition, with other MSC-2 members being central in carrying out later steps. These results illustrate how multi-omics applied to a defined consortium untangles interactions between soil microorganisms.

McClure, Ryan [Pacific Northwest National Laborato↗

Gut microbiota carbon and sulfur metabolisms support Salmonella infections

Abstract Salmonella enterica serovar Typhimurium is a pervasive enteric pathogen and ongoing global threat to public health. Ecological studies in the Salmonella impacted gut remain underrepresented in the literature, discounting microbiome mediated interactions that may inform Salmonella physiology during colonization and infection. To understand the microbial ecology of Salmonella remodeling of the gut microbiome, we performed multi-omics on fecal microbial communities from untreated and Salmonella-infected mice. Reconstructed genomes recruited metatranscriptomic and metabolomic data providing a strain-resolved view of the expressed metabolisms of the microbiome during Salmonella infection. These data informed possible Salmonella interactions with members of the gut microbiome that were previously uncharacterized. Salmonella-induced inflammation significantly reduced the diversity of genomes that recruited transcripts in the gut microbiome, yet increased transcript mapping was observed for seven members, among which Luxibacter and Ligilactobacillus transcript read recruitment was most prevalent. Metatranscriptomic insights from Salmonella and other persistent taxa in the inflamed microbiome further expounded the necessity for oxidative tolerance mechanisms to endure the host inflammatory responses to infection. In the inflamed gut lactate was a key metabolite, with microbiota production and consumption reported amongst members with detected transcript recruitment. We also showed that organic sulfur sources could be converted by gut microbiota to yield inorganic sulfur pools that become oxidized in the inflamed gut, resulting in thiosulfate and tetrathionate that support Salmonella respiration. This research advances physiological microbiome insights beyond prior amplicon-based approaches, with the transcriptionally active organismal and metabolic pathways outlined here offering intriguing intervention targets in the Salmonella-infected intestine.

59 BASIC BIOLOGICAL SCIENCES↗

Temporal dynamics of the multi-omic response to endurance exercise training

Regular exercise promotes whole-body health and prevents disease, but the underlying molecular mechanisms are incompletely understood. Here, the Molecular Transducers of Physical Activity Consortium profiled the temporal transcriptome, proteome, metabolome, lipidome, phosphoproteome, acetylproteome, ubiquitylproteome, epigenome and immunome in whole blood, plasma and 18 solid tissues in male and female Rattus norvegicus over eight weeks of endurance exercise training. The resulting data compendium encompasses 9,466 assays across 19 tissues, 25 molecular platforms and 4 training time points. Thousands of shared and tissue-specific molecular alterations were identified, with sex differences found in multiple tissues. Temporal multi-omic and multi-tissue analyses revealed expansive biological insights into the adaptive responses to endurance training, including widespread regulation of immune, metabolic, stress response and mitochondrial pathways. Many changes were relevant to human health, including non-alcoholic fatty liver disease, inflammatory bowel disease, cardiovascular health and tissue injury and recovery. The data and analyses presented in this study will serve as valuable resources for understanding and exploring the multi-tissue molecular effects of endurance training and are provided in a public repository (https://motrpac-data.org/).

59 BASIC BIOLOGICAL SCIENCES↗

Using supervised machine-learning approaches to understand abiotic stress tolerance and design resilient crops

Abiotic stresses such as drought, heat, cold, salinity and flooding significantly impact plant growth, development and productivity. As the planet has warmed, these abiotic stresses have increased in frequency and intensity, affecting the global food supply and making it imperative to develop stress-resilient crops. In the past 20 years, the development of omics technologies has contributed to the growth of datasets for plants grown under a wide range of abiotic environments. Integration of these rapidly growing data using machine-learning (ML) approaches can complement existing breeding efforts by providing insights into the mechanisms underlying plant responses to stressful conditions, which can be used to guide the design of resilient crops. In this review, we introduce ML approaches and provide examples of how researchers use these approaches to predict molecular activities, gene functions and genotype responses under stressful conditions. Finally, we consider the potential and challenges of using such approaches to enable the design of crops that are better suited to a changing environment. This article is part of the theme issue ‘Crops under stress: can we mitigate the impacts of climate change on agriculture and launch the ‘Resilience Revolution’?’.

abiotic stress↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology↗

1 × 1 km maps of abundances of eight enzyme functional classes for soil C, N, and P cycling across the CONUS

This dataset includes eight 1 × 1 km maps of the abundances of eight enzyme functional classes (EFC) for soil C, N, and P cycling across the CONUS. These mappings are predicted by the machine learning model trained using metagenomics and the corresponding environmental data. This item corresponds to our article: Fan, C., Song, Y., Mishra, U., Gautam, S., & Mayes, M. A. (2025). Harnessing the Power of Machine Learning and Omics to Identify Environmental Regulation on Microbial Functional Composition for Soil C, N, and P Cycling. Journal of Geophysical Research: Biogeosciences, 130(10).

1 × 1 km↗

Harnessing the Power of Machine Learning and Omics to Identify Environmental Regulation on Microbial Functional Composition for Soil C, N, and P Cycling

Microbial enzyme-mediated soil organic matter (SOM) decomposition regulates many key ecosystem functions, such as elemental cycling, soil carbon sequestration, and soil fertility. However, representing microbial processes in Earth system models (ESMs) remains challenging due to a limited understanding of the spatial patterns of diverse microbial functions responsible for soil carbon (C), nitrogen (N), and phosphorus (P) cycling as well as the underlying mechanisms regulating their relative abundances across various environments. We collected published metagenomics data across the continental US (CONUS) to identify hundreds of microbial genes involved in soil C, N, and P cycling and grouped them into eight enzyme functional classes (EFCs). Each EFC represented a group of gene-encoded potential enzymes that decompose similar soil compounds. By integrating the abundances of omics-informed EFCs with the corresponding environmental information, we trained a machine learning (ML) model to identify key edaphic, climate, and vegetation factors regulating the abundances of each EFC. Quantitative analysis of effects of these factors revealed that the spatial distribution of eight EFCs for soil C, N, and P cycling across CONUS reflected potential resource optimization strategies of microbial communities under nutrient limitation, preferential organic-mineral associations, and climatological stresses. This insight, together with the interpreted ML tool and the CONUS-level benchmark for EFCs abundances, paves the way for parameterizing environmental-regulated microbial functional dynamics in biogeochemical models.

machine learning↗

Altering translation allows E. coli to overcome G-quadruplex stabilizers

The data included in this Dryad submission was collected in order to understand how the model organism* Escherichia coli* overcomes stabilized G-quadruplexes. This work involved a multi-omics approach to studying how the G-quadruplex stabilizers NMM and Braco-19 impact growth, gene importance, and mRNA/proteomic abundance in G-quadruplex stabilizing conditions.

Bacteria↗