Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

High-throughput single-cell transcriptomics of bacteria using combinatorial barcoding

Microbial split-pool ligation transcriptomics (microSPLiT) is a high-throughput single-cell RNA sequencing method for bacteria. With four combinatorial barcoding rounds, microSPLiT can profile transcriptional states in hundreds of thousands of Gram-negative and Gram-positive bacteria in a single experiment without specialized equipment. As bacterial samples are fixed and permeabilized before barcoding, they can be collected and stored ahead of time. During the first barcoding round, the fixed and permeabilized bacteria are distributed into a 96-well plate, where their transcripts are reverse transcribed into cDNA and labeled with the first well-specific barcode inside the cells. The cells are mixed and redistributed two more times into new 96-well plates, where the second and third barcodes are appended to the cDNA via in-cell ligation reactions. Finally, the cells are mixed and divided into aliquot sub-libraries, which can be stored until future use or prepared for sequencing with the addition of a fourth barcode. It takes 4 days to generate sequencing-ready libraries, including 1 day for collection and overnight fixation of samples. Here, the standard plate setup enables single-cell transcriptional profiling of up to 1 million bacterial cells and up to 96 samples in a single barcoding experiment, with the possibility of expansion by adding barcoding rounds. The protocol requires experience in basic molecular biology techniques, handling of bacterial samples and preparation of DNA libraries for next-generation sequencing. It can be performed by experienced undergraduate or graduate students. Data analysis requires access to computing resources, familiarity with Unix command line and basic experience with Python or R.

59 BASIC BIOLOGICAL SCIENCES↗

Legacy Effects of Cropping System and Precipitation Influence the Core Camelina sativa Microbiome

Camelina ( Camelina sativa L.) is a potential biofuel crop and beneficial rotation crop in dryland cropping systems. Little is known about camelina microbiota or the legacy effect of soil origin/cropping system zones on camelina-associated microbiome assembly. To explore camelina-microbe associations, we grew camelina in the greenhouse using soil transplanted from 33 locations in the dryland wheat production area of eastern Washington. Bacterial, archaeal, and fungal communities from bulk soil, rhizosphere, and endosphere were characterized with 16S rRNA and internal transcribed spacer amplicon sequencing and were analyzed alongside site-specific climatic and edaphic data. We found that soil from the highest precipitation zone had higher alpha diversity than soil from the driest zone, but this effect was not seen in the greenhouse rhizosphere or endosphere. Plant compartment, cropping system zone, and soil origin all significantly influenced microbial composition, with soil pH and organic matter, as well as precipitation at origin, as major predictors. Analysis of abundance–occupancy distributions showed that the Actinobacteriota Aeromicrobium and Marmoricola and the fungus Pseudogymnoascus in the rhizosphere were plant-selected, while the endosphere was characterized by a number of Actinobacteriota, Rhizobium, and Clostridium. Sphingomonas amplicon sequence variants were also consistently enriched in the rhizosphere, suggesting that they are present in soils collected throughout eastern Washington and may represent good candidate biostimulants. Several lignin decomposing fungi had site-specific rhizospheric distributions, suggesting that they may be dispersal-limited or result from the legacy effect of long-term wheat cropping. Overall, this study contributes to our understanding of microbiome assembly in and on camelina roots while also highlighting the potential impact of cropping history on soil- and plant-associated microbiomes. [Formula: see text] The author(s) have dedicated the work to the public domain under the Creative Commons CC0 “No Rights Reserved” license by waiving all of his or her rights to the work worldwide under copyright law, including all related and neighboring rights, to the extent allowed by law, 2025.

Barnes, Elle M↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Predicting river turbidity in Pine Island Bayou using machine learning techniques coupled with variational mode decomposition

Elevated turbidity levels pose significant public health risks by facilitating the transport of harmful pollutants, including metals, organic compounds, and pathogenic microorganisms into the surface water. These conditions create serious challenges for public recreational water use and drinking water treatment, leading to economic losses and health risks. This study utilizes water monitoring data in Pine Island Bayou, Texas, and develops a Sequence-to-Sequence (S2S) model to predict turbidity using Attention-based Gated Recurrent Units with Encoder-Decoder (AT-GRU-ED) and Long Short-Term Memory (LSTM), coupled with Variational Mode Decomposition (VMD). Compared to the model without VMD, the model demonstrates satisfactory 72-hour turbidity prediction performance, achieving MAEs of 2.60 and 3.29 NTU (reductions of 53% and 58%), RMSEs of 21.08 and 31.49 NTU (reductions of 82% and 80%), and R² values of 0.96 and 0.84 on the validation and test sets, respectively. Feature importance analysis reveals that water temperature is the dominant factor influencing seasonal turbidity patterns, while real-time hourly rainfall significantly contributes to short-term variability. Turbidity typically peaks within 48 hours after rainfall events due to lagged effects from surface runoff and upstream flow. Findings suggest suspending recreational water use and water supply pumping for three days after heavy rainfall can benefit public health and improve water treatment processes. Discharges above 100 m3/s are found to accelerate sediment dilution and transport, reducing turbidity levels more quickly after the peak. In conclusion, the proposed model demonstrates reliable 72-hour turbidity prediction, supporting decision-making for water treatment plant operations and providing early warning for public recreational water use.

Deep learning↗

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

Human Liver Epithelial Cells (HuH7) Response to HCoV-229E Infection Epigenomics (ATAC-Seq) (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-299E) infection alters chromatin accessibility in infected cells. Sample data was obtained from mock-infected cells, UV-inactivated virus treated cells, and replication competent HCoV-229E infected immortalized human liver cells (HuH7) at 24 hours post infection. Samples were processed using ATAC-seq methods for reported bar coded libraries. Sample data was acquired using an Illumina Hi-Seq 2500 sequencer system and further processed for ATAC-Seq expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Human Primary Airway Epithelium +/- Macrophages Response to HCoV-229E Infection Transcriptomics (ACS-DP3)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-299E) infection. Sample data was obtained for mock and infected (MOI 3) primary human airway epithelial cells with and without macrophages and grown in air-liquid interface conditions. Sample data was acquired using an Illumina Hi-Seq 4000 sequencer system and further processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular and Epidemiological Investigation of Fluconazole-resistant Candida parapsilosis —Georgia, United States, 2021

Abstract Background Reports of fluconazole-resistant Candida parapsilosis bloodstream infections are increasing. We describe a cluster of fluconazole-resistant C parapsilosis bloodstream infections identified in 2021 on routine surveillance by the Georgia Emerging Infections Program in conjunction with the Centers for Disease Control and Prevention. Methods Whole-genome sequencing was used to analyze C parapsilosis bloodstream infections isolates. Epidemiological data were obtained from medical records. A social network analysis was conducted using Georgia Hospital Discharge Data. Results Twenty fluconazole-resistant isolates were identified in 2021, representing the largest proportion (34%) of fluconazole-resistant C parapsilosis bloodstream infections identified in Georgia since surveillance began in 2008. All resistant isolates were closely genetically related and contained the Y132F mutation in the ERG11 gene. Patients with fluconazole-resistant isolates were more likely to have resided at long-term acute care hospitals compared with patients with susceptible isolates (P = .01). There was a trend toward increased mechanical ventilation and prior azole use in patients with fluconazole-resistant isolates. Social network analysis revealed that patients with fluconazole-resistant isolates interfaced with a distinct set of healthcare facilities centered around 2 long-term acute care hospitals compared with patients with susceptible isolates. Conclusions Whole-genome sequencing results showing that fluconazole-resistant C parapsilosis isolates from Georgia surveillance demonstrated low genetic diversity compared with susceptible isolates and their association with a facility network centered around 2 long-term acute care hospitals suggests clonal spread of fluconazole-resistant C parapsilosis. Further studies are needed to better understand the sudden emergence and transmission of fluconazole-resistant C parapsilosis.

Misas, Elizabeth (ORCID:0000000162437716)↗

Solar fusion III: New data and theory for hydrogen-burning stars

In stars that lie on the main sequence in the Hertzsprung-Russell diagram, like our Sun, hydrogen is fused to helium in a number of nuclear reaction chains and series, such as the proton-proton chain and the carbon-nitrogen-oxygen cycles. Precisely determined thermonuclear rates of these reactions lie at the foundation of the standard solar model. This review, the third decadal evaluation of the nuclear physics of hydrogen-burning stars, is motivated by the great advances made in recent years by solar neutrino observatories, putting experimental knowledge of the proton-proton (𝑝⁢𝑝)-chain neutrino fluxes in the few-percent precision range. The basis of the review is a one-week community meeting held in July 2022 in Berkeley, California, and many subsequent digital meetings and exchanges. The relevant reactions of solar and stellar hydrogen burning are reviewed here from both theoretical and experimental perspectives. Recommendations for the state of the art of the astrophysical 𝑆 factor and its uncertainty are formulated for each of them. Furthermore, several other topics of paramount importance for the solar model are reviewed as well: recent and future neutrino experiments, electron screening, radiative opacities, and current and upcoming experimental facilities. In addition to reaction-specific recommendations, general recommendations are also formed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Prioritizing Uncertainties in Hydrogen Contribution to Risk in Post-Crash Outcomes for Rail

This report presents analysis from Sandia National Laboratories predicting contributions to risk associated with the use of hydrogen technology for rail. Event sequence diagrams are used to describe possible accident scenarios and progressions. Initiating event frequencies and branch event probabilities for each scenario are quantified with uncertainty using distributions fit to Federal Railroad Administration and U.S. Department of Transportation Pipeline and Hazardous Materials Safety Administration data on applicable accidents from 2000 to 2020. Uncertainty is propagated through the event sequence diagram to estimate the frequency and conditional probability of accident end states. The analysis identifies four scenarios with significant contributions to risk from hydrogen that are predicted to occur relatively frequently, which may inform priorities for reducing uncertainty. These scenarios are 1) overpressure events resulting from collisions with hydrogen release due to mechanical damage and delayed ignition, 2) jet fire events resulting from collisions with hydrogen release due to mechanical damage and immediate ignition, 3) jet fires resulting from fire or explosion initiating events involving the hydrogen tank and correct operation of the thermally-activated pressure relief device (TPRD) subsequent to the thermal insult, and 4) pressure burst resulting from fire or explosion initiating events involving the hydrogen tank and failure of the TPRD. Delayed and immediate hydrogen ignition probabilities are identified as being highly uncertain and potential candidates for reducing conservatism in the predicted frequencies for these two scenarios.

08 HYDROGEN↗

Direct local parametrization of nuclear state densities using the back-shifted Bethe formula

Level densities are often parametrized using the back-shifted Bethe formula (BBF) for nuclei that possess experimental data for s-wave neutron resonance average spacings and a complete discrete level sequence at low excitation energies. However, these parametrizations require the additional modeling of the dependence of the spin-cutoff parameter on excitation energy. Here, in this work, we avoid the need to model the spin distribution of level densities by using the experimental data to parametrize directly the state densities, for which the BBF does not depend on the spin-cutoff parameter. This approach allows for a local parameterization of state densities that is independent of the spin-cutoff parameter. We provide these parameters in a tabulated form for applications in nuclear reaction calculations and for testing microscopic approaches to state densities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

KBase Narrative - Complete genome sequence of a novel Microbacterium sp. strain Clip185.

We have isolated a new species of Microbacterium, an Actinobacterium. We have temporarily named this bacterium as Microbacterium sp. strain Clip185 (hereafter called strain Clip185) from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii strain LMJ.RY0402.185141 (a Chlamydomonas Library project CLiP strain). We sequenced the whole genome of strain Clip185 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of this new Microbacterium species that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

Mitra, Mautusi↗

Complete genome sequence of Sphingobium yanoikuyae strain CC4533

We have isolated a new strain of Sphingobium yanoikuyae , which belongs to the class Alphaproteobacteria, order Sphingomonadales, and family Sphingomonadaceae. This carotenoid-producing strain is capable of degrading xenobiotics and is tolerant to toxic levels of six heavy metals. We have designated the newly isolated strain of S. yanoikuyae as S. yanoikuyae strain CC4533 (hereafter called strain CC4533) because it was isolated from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii wild type strain CC4533. We sequenced the whole genome of strain CC4533 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of S. yanoikuyae strain CC4533 that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset_for_Molecular_Motion_Below_the_Glass_Transition_A_Solid-State_NMR_Study_of_Siloxane_Polymer_Dynamics Study

This dataset contains solid-state 1H and 13C NMR relaxometry data, differential scanning calorimetry (DSC) data, and size exclusion chromatography (SEC/GPC) data supporting the study of sub-glass-transition (sub-Tg) molecular dynamics in a composition- and sequence-controlled series of diphenyl-substituted polysiloxanes (PDMS, 14Ph, 33Ph, 50Ph, 67Ph, and 100Ph; 0–100% diphenylsiloxane content by mole).All solid-state NMR data were acquired on a 200 MHz Bruker Avance III HD spectrometer using a static 7 mm HX probe or a 4 mm HX probe under 4 kHz magic-angle spinning. Raw Bruker TopSpin experiment folders are included for: (1) variable-temperature 1H lineshape measurements used to determine linewidth (FWHM) as a function of temperature across the glass transition; (2) 1H T1 (saturation recovery with solid-echo detection), probing nanosecond-scale dynamics near the 1H Larmor frequency; (3) 1H T1rho (direct spin-lock, 62.5 kHz), probing microsecond-scale segmental dynamics; (4) 13C-detected Lee–Goldburg cross-polarization 1H T1rho (LGCPH T1rho) for 33Ph and 50Ph, resolving aromatic and aliphatic proton environments; and (5) 13C T1 relaxation for 33Ph and 50Ph. Differential scanning calorimetry data (TA Instruments DSC 25, −150 to +120 °C, up to +300 °C for 100Ph, 10 °C/min) are included for all six compositions and support the glass-transition temperatures in Table 1 and Figure 1. Size exclusion chromatography data (Agilent 1200 Series, PL-Gel 300 mixed-C column, THF mobile phase, polystyrene calibration standards) are included for the three synthesized copolymers (33Ph, 50Ph, 67Ph) and support the number-average molecular weights in Table 1. Processed data include per-composition relaxation-time summaries (Excel), curve-fitting and Bloembergen-Purcell-Pound (BPP) model analysis notebooks (Jupyter/Python), and Igor Pro (.pxp) master files used to generate the manuscript's figures.

Bloembergen-Purcell-Pound theory↗

The SRG/eROSITA All-Sky Survey: Optical identification and properties of galaxy clusters and groups in the western galactic hemisphere

The first SRG/eROSITA All-Sky Survey (eRASS1) provides the largest intracluster medium-selected galaxy cluster and group catalog covering the western Galactic hemisphere. Compared to samples selected purely on X-ray extent, the sample purity can be enhanced by identifying cluster candidates using optical and near-infrared data from the DESI Legacy Imaging Surveys. Using the red-sequence-based cluster findereROMaPPer, we measured individual photometric properties (redshiftz λ , richnessλ, optical center, and BCG position) for 12000 eRASS1 clusters over a sky area of 13 116 deg 2 , augmented by 247 cases identified by matching the candidates with known clusters from the literature. The median redshift of the identified eRASS1 sample isz= 0.31, with 10% of the clusters atz> 0.72. The photometric redshifts have an accuracy ofδz/(1 +z) ≲ 0.005 for 0.05 specand velocity dispersionσ) were measured a posteriori for a subsample of 3210 and 1499 eRASS1 clusters, respectively, using an extensive compilation of spectroscopic redshifts of galaxies from the literature. We infer that the primary eRASS1 sample has a purity of 86% and optical completeness >95% forz> 0.05. For these and further quality assessments of the eRASS1 identified catalog, we applied our identification method to a collection of galaxy cluster catalogs in the literature, as well as blindly on the full Legacy Surveys covering 24069 deg 2 . Using a combination of these cluster samples, we investigated the velocity dispersion-richness relation, finding that it scales with richness as log(λ norm ) = 2.401 × log(σ) − 5.074 with an intrinsic scatter ofδ in = 0.10 ± 0.01 dex. The primary product of our work is the identified eRASS1 cluster catalog with high purity and a well-defined X-ray selection process, opening the path for precise cosmological analyses presented in companion papers.

Astronomy & Astrophysics↗

Evaluation of DNA Extraction Efficiency in Diverse Algae Strains Using Commercial Kits and Lysis Approaches

Efficient DNA extraction is essential for accurately monitoring microalgae communities in large-scale cultivation systems such as raceway ponds and wastewater ponds. Traditional phenol chloroform extracts are a staple in microbiology but are obsolete for routine sampling due to its high toxicity reagents and time intensive setups. Commercial DNA extraction kits are more favorable for the microbes found in these ponds, but lack specific kits made for these communities. Little is known about which kits perform the best, leading researchers to use a variety of different kits with inconsistent results. This project compared one precipitation based commercial kit (Lucigen Masterpure) and five wash based kits (Monarch, Zymo Quick-DNA, and three Qiagen DNeasy kits) using four brackish algae strains to determine which methods yield the greatest quantity and quality of genomic DNA. Extractions were evaluated using the manufacturers protocol, and additional pretreatment options were administered before a single kit to compare its potential in being added routinely before extractions. Pretreatment options included both cryogenic freeze-thawing and heat incubation using enzymes. DNA was quantified using Qubit fluorometry and NanoDrop purity ratios. Overall, the Qiagen PowerWater kit provided the highest DNA yield and purity, but at a significantly higher cost then the precipitation-based kit (MasterPure). It was also noted that while the precipitation-based kit was significantly cheaper, provided similar results, it took significantly more time to complete a single run. Cryogenic pretreatment (6x cycles) increased average DNA yields by up to 80%, whereas enzymatic pretreatment most improved purity ratios without substantially improving quantity. The results suggest that it may be more cost and time efficient to use Qiagen kits with the addition of lysis pretreatments to procure better results. Future works includes developing a better system to efficiently collect multi variable data, and to upscale to artificial polycultures using similar methodologies alongside sequencing to confirm kit results.

59 BASIC BIOLOGICAL SCIENCES↗