Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

Multiple-amplifier sensing charged-coupled device: model and improvement of the node removal efficiency

The multiple-amplifier sensing charge-coupled device (MAS-CCD) has emerged as a promising technology for astronomical observation, quantum imaging, and low-energy particle detection due to its ability to reduce the readout time for the same readout noise level compared with its predecessor, the skipper-CCD, by reading out the same charge packet through multiple inline amplifiers. Previous works identified a new parameter in this sensor, called node removal inefficiency (NRI), related to inefficiencies in charge transfer and residual charge removal from the sense node of each amplifier after readout. These inefficiencies can lead to distortions in the measured signals similar to those produced by the charge transfer inefficiencies in standard CCDs. We introduce more details in the mathematical model of the NRI mechanism and provide techniques to quantify its magnitude from the measured data. It also proposes a new operation strategy that significantly reduces its effect with minimal alterations of the timing sequences or voltage settings for the other signals of the sensor. The proposed technique is demonstrated experimentally on a 16-amplifier MAS-CCD. At the same time, the experimental data demonstrate that this approach minimizes the NRI effect to levels comparable with other sources of distortion such as the charge transfer inefficiency in scientific devices.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Apatite geochemistry as a tool for understanding the petrogenesis of layered mafic-ultramafic rocks in the Bushveld Complex, South Africa

The sources of the magmas that formed the Rustenburg Layered Suite of the Bushveld Complex in South Africa remain debated, despite decades of research. Vertical and lateral variation in bulk rock and mineral separate Sr-Nd isotopic compositions, which generally indicate enriched sources, demonstrate that the layered sequence was formed by the emplacement of multiple batches of magma, crucially resulting in episodes of PGE-Cr-V mineralisation. The Lu-Hf isotope compositions of zircon are, however, at odds with the bulk rock Sr-Nd isotopic heterogeneity as they show near homogeneous compositions throughout the layered sequence (εHf (2.06 Ga) =−8). This lack of variation in Hf isotope composition has been attributed to deep, continental lithospheric mantle-related and/or crustal contamination of plume-derived Bushveld magmas. In this study, we analysed the major, trace element and Sr-Nd isotope geochemistry of apatite in the Rustenburg Layered Suite. Apatite occurs as an intercumulus mineral in the lowermost regions and a cumulus mineral in the uppermost regions of the layered sequence and can therefore be used to test existing models for the isotopic disequilibrium between bulk rock Sr-Nd and zircon Hf isotopic compositions. Apatite is largely chlorapatite in the lowermost regions and fluorapatite in the uppermost regions of the layered sequence. The Merensky Reef is unusual in that it comprises both chlorapatite and fluorapatite. Apatite throughout the layered sequence is generally unzoned and shows no evidence of late-stage alteration. Trace element data show that apatite is enriched in L/HREE, with common negative Eu-Sr anomalies. These trace element signatures are consistent with a magmatic origin for the apatite grains, with prior, or concurrent, plagioclase crystallization from the same melt. Variability in in situ Sr and Nd isotope compositions of apatite is recorded throughout the layered sequence with εNd (2.06 Ga) compositions varying between −2.5 and − 10.2 and initial 87 Sr/ 86 Sr compositions varying between 0.7079 and 0.7103 (for the Marikana dikes only). The variability in Sr-Nd isotope compositions of apatite is consistent with the bulk rock (and mineral separate) variation in Sr-Nd isotope compositions, suggesting apatite preserves primary magmatic compositions in the Rustenburg Layered Suite.

Apatite↗

Protocol for applying a network-enabled gene discovery pipeline to non-model plant species

Identifying upstream regulators of key genes is essential for understanding gene regulatory mechanisms and translating these insights into functional targets. Here, we present a protocol for applying the network-enabled gene discovery pipeline (NEEDLE) to non-model plant species. We describe steps for environment setup, data preparation, computational analysis, expected outputs, and parameter considerations. NEEDLE integrates RNA sequencing (RNA-seq) processing, weighted gene co-expression analysis (WGCNA), Gene Network Inference with Ensemble of trees (GENIE3), and promoter conservation analysis to prioritize candidate transcriptional regulators.

Plant Sciences↗

Host population dynamics influence Leptospira spp. transmission patterns among Rattus norvegicus in Boston, Massachusetts, US

Leptospirosis (caused by pathogenic bacteria in the genus Leptospira ) is prevalent worldwide but more common in tropical and subtropical regions. Transmission can occur following direct exposure to infected urine from reservoir hosts, or a urine-contaminated environment, which then can serve as an infection source for additional rats and other mammals, including humans. The brown rat, Rattus norvegicus , is an important reservoir of Leptospira spp. in urban settings. We investigated the presence of Leptospira spp. among brown rats in Boston, Massachusetts and hypothesized that rat population dynamics in this urban setting influence the transportation, persistence, and diversity of Leptospira spp. We analyzed DNA from 328 rat kidney samples collected from 17 sites in Boston over a seven-year period (2016–2022); 59 rats representing 12 of 17 sites were positive for Leptospira spp. We used 21 neutral microsatellite loci to genotype 311 rats and utilized the resulting data to investigate genetic connectivity among sampling sites. We generated whole genome sequences for 28 Leptospira spp. isolates obtained from frozen and fresh tissue from some of the 59 positive rat kidneys. When isolates were not obtained, we attempted genomic DNA capture and enrichment, which yielded 14 additional Leptospira spp. genomes from rats. We also generated an enriched Leptospira spp. genome from a 2018 human case in Boston. We found evidence of high genetic structure among rat populations that is likely influenced by major roads and/or other dispersal barriers, resulting in distinct rat population groups within the city; at certain sites these groups persisted for multiple years. We identified multiple distinct phylogenetic clades of L. interrogans among rats that were tightly linked to distinct rat populations. This pattern suggests L. interrogans persists in local rat populations and its transportation is influenced by rat population dynamics. Finally, our genomic analyses of the Leptospira spp. detected in the 2018 human leptospirosis case in Boston suggests a link to rats as the source. These findings will be useful for guiding rat control and human leptospirosis mitigation efforts in this and other similar urban settings.

Stone, Nathan E.↗

Citywide indoor air sampling mirrors wastewater and clinical for environmental surveillance of respiratory viruses

Wastewater surveillance of respiratory pathogens can provide timely estimates of viral activity and disease trends in a population. Indoor air surveillance could be used similarly with some advantages but remains largely unvalidated at the community -scale. Here, an indoor air surveillance program was employed as part of public health environmental surveillance in Chicago, Illinois, USA. Ten air samplers were placed in healthcare and congregate living settings across the city. Weekly air samples were evaluated for influenza A, influenza B, respiratory syncytial virus, and SARS -CoV-2 over two respiratory virus seasons (2023 -2025). Citywide, aggregated air sample positivity and viral load were closely correlated with local clinical case and wastewater surveillance data across all respiratory viruses. Virus trends in air data often preceded clinical and wastewater, although this varied across pathogens and respiratory virus seasons. Further, whole -genome sequencing of SARS -CoV-2 showed close correlation of variant proportions across all datasets. At the building -scale, air samples obtained from a single sampling device provided efficient respiratory virus surveillance, with respiratory pathogen levels mirroring citywide clinical surveillance data. These data demonstrate that air surveillance can provide respiratory virus case and variant trend data at a building or community -scale, serving as an alternative or complementary tool for public health environmental surveillance.

Wilton, Rosemarie↗

RB-TnSeq elucidates dicarboxylic-acid-specific catabolism in β-proteobacteria for improved plastic monomer upcycling

Dicarboxylic acids are key components of many polymers and plastics, making them a target for both engineered microbial degradation and sustainable bioproduction. In this study, we generated a comprehensive data set of functional evidence for the genetic basis of dicarboxylic and fatty acid metabolism using randomly barcoded transposon sequencing (RB-TnSeq). We identified four β-proteobacteria that displayed robust growth with dicarboxylic acid sole carbon source and cultured their mutant libraries with dicarboxylic and fatty acids with carbon chain lengths from C3 to C12. The resulting fitness data suggested that dicarboxylic and fatty acid metabolisms are largely distinct, and different sets of β-oxidation genes are required for catabolizing dicarboxylic versus fatty acids of the same carbon chain lengths. In addition, we identified transcriptional regulators and transporters with strong fitness phenotypes related to dicarboxylic acid utilization. In Ralstonia sp. UNC404CL21Col (R. CL21), we deleted two transcriptional repressors to improve its utilization of short-chain dicarboxylic acids. We exploited the diacid-utilizing catabolism of R. CL21 to upcycle a mock mixture of the dicarboxylic acids produced when polyethylene is oxidized. After introducing a heterologous indigoidine production pathway, this engineered Ralstonia produced 0.56 ± 0.02 g/L indigoidine from a mixture of dicarboxylic acids as a carbon source, demonstrating the potential of R. CL21 to upcycle plastic wastes to products derived from tricarboxylic acid (TCA) cycle intermediates. IMPORTANCE: Upcycling the carbon in plastic wastes to value-added products is a promising approach to address the plastic waste and climate crises, and dicarboxylic acid metabolism is an important facet of several approaches. Improving our understanding of the genetic basis of this metabolism has the potential to uncover new enzymes and genetic parts for engineered pathways involving dicarboxylic acids. Our data set is the most comprehensive interrogation of dicarboxylic acid catabolism to date, and this work will be of utility to researchers interested in both plastics bioproduction and upcycling applications.

Pearson, Allison N↗

Recurrent features of amplitudes in planar $\mathcal{N}$ = 4 super Yang-Mills theory

The planar three-gluon form factor for the chiral stress tensor operator in planar maximally supersymmetric Yang-Mills theory is an analog of the Higgs-to-three-gluon scattering amplitude in QCD. The amplitude (symbol) bootstrap program has provided a wealth of high-loop perturbative data about this form factor, with results up to eight loops available. The symbol of the form factor at L loops is given by words of length 2L in six letters with associated integer coefficients. In this paper, we analyze this data, describing patterns of zero coefficients and relations between coefficients. We find many sequences of words whose coefficients are given by closed-form expressions which we expect to be valid at any loop order. Moreover, motivated by our previous machine-learning analysis, we identify simple recursion relations that relate the coefficient of a word to the coefficients of particular lower-loop words. These results open an exciting door for understanding scattering amplitudes at all loop orders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Tuning Copolymer Microstructure Using Ring-Opening Cross-Metathesis Polymerization

The capability of ring-opening cross-metathesis (RO/CM) polymerization to produce alternating copolymers was studied. By treating commercial polybutadiene (PB) with bulky oxanorbornene monomers and Ru-based olefin metathesis catalysts, alternating copolymers were produced under mild conditions with high sequence fidelities. Here, we found that alternating copolymers could be produced starting from a variety of butadiene sources including PB, cyclooctadiene (COD), or t,t,t-1,5,9-cyclododecatriene (CDT), highlighting for the first time the kinetic pathway independence of this process. Kinetic copolymerization analysis of an oxanorbonene monomer with CDT revealed that much higher monomer conversions were obtained compared with the analogous homopolymerizations and showed evidence of alternating monomer incorporation. Copolymerization of these monomers also enabled good control when targeting different molecular weights. Copolymer thermal analysis revealed a strong correlation between thermal behavior and alternating sequence fidelity, providing a second lever beyond composition to tune thermal behavior. These data demonstrate that a broad variety of polymer microstructures can be accessed via RO/CM polymerization and highlight the potential of CDT in alternating copolymer synthesis.

Foster, Jeffrey C. [Oak Ridge National Laboratory ↗

Unveiling a pervasive DNA adenine methylation regulatory network in the early-diverging fungus Rhizopus microsporus

Development of the DNA affinity purification and sequencing (DAP-seq) technique has allowed genome-scale studies of transcription factor (TF)-binding sites with high reproducibility. Here, we apply this technique to the human opportunistic pathogen Rhizopus microsporus, a mucoralean fungus belonging to the understudied group of early-diverging fungi. We characterize genome-wide binding sites of 58 TFs encoded by genes regulated through adenine methylation and representing major TF families. This analysis reveals their binding profiles and recognized sequences, expanding and diversifying the catalog of known fungal motifs. By integrating this data with DNA 6-methyladenine profiling, we uncover the extensive direct and indirect impact of this epigenetic modification on the regulation of gene expression. Furthermore, we use the generated data to identify TFs involved in biologically relevant processes such as zinc metabolism and light response. Our work enhances our understanding of regulatory mechanisms in R. microsporus and provides broader insights into gene regulation across the fungal kingdom.

Lax, Carlos [Universidad de Murcia (Spain)] (ORCID↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Combinatorial transcription factor binding encodes cis -regulatory wiring of mouse forebrain GABAergic neurogenesis

Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.

59 BASIC BIOLOGICAL SCIENCES↗

Atmospheric parameters and chemical abundances within 100 pc: a sample of G, K, and M main-sequence stars

ABSTRACT To date, we have access to enormous inventories of stellar spectra that allow the extraction of atmospheric parameters and chemical abundances essential in stellar studies. However, characterizing such a large amount of data is complex and requires a good understanding of the studied object to ensure reliable and homogeneous results. In this study, we present a methodology to measure homogenously the basic atmospheric parameters and detailed chemical abundances of over 1600 thin disc main-sequence stars in the 100 pc solar neighbourhood, using APOGEE-2 infrared spectra. We employed the code tonalli to determine the atmospheric parameters using a prior on $\log {g}$. The $\log {g}$ prior in tonalli implies an understanding of the treated population and helps to find physically coherent answers. Our atmospheric parameters agree within the typical uncertainties (100 K in $\mathrm{T_{eff}}$, 0.15 dex in $\log {g}$ and [M/H]) with previous estimations of ASPCAP and Gaia DR3. We use our temperatures to determine a new infrared colour–temperature sequence, in good agreement with previous works, that can be used for any main-sequence star. Additionally, we used the bacchus code to determine the abundances of Mg, Al, Si, Ca, and Fe in our sample. The five elements (Mg, Al, Si, Ca, Fe) studied have an abundance distribution centred around slightly subsolar values in agreement with previous results for the solar neighbourhood. The over 1600 main-sequence stars’ atmospheric parameters and chemical abundances presented here are useful in follow-up studies of the solar neighbourhood or as a training set for data-driven methods.

López-Valdivia, Ricardo (ORCID:0000000277950018)↗

Deep learning with plasma plume image sequences for anomaly detection and prediction of growth kinetics during pulsed laser deposition

Abstract Materials synthesis platforms that are designed for autonomous experimentation are capable of collecting multimodal diagnostic data that can be utilized for feedback to optimize material properties. Pulsed laser deposition (PLD) is emerging as a viable autonomous synthesis tool, and so the need arises to develop machine learning (ML) techniques that are capable of extracting information from in situ diagnostics. Here, we demonstrate that intensified-CCD image sequences of the plasma plume generated during PLD can be used for anomaly detection and the prediction of thin film growth kinetics. We develop multi-output (2 + 1)D convolutional neural network regression models that extract deep features from plume dynamics that not only correlate with the measured chamber pressure and incident laser energy, but more importantly, predict parameters of an auto-catalytic film growth model derived from in situ laser reflectivity experiments. Our results demonstrate how ML with in situ plume diagnostics data in PLD can be utilized to maintain deposition conditions in an optimal regime. Further, the predictive capabilities of plume dynamics on the kinetics of film growth or other film properties prior to deposition provides a means for rapid pre-screening of growth conditions for the non-expert, which promises to accelerate materials optimization with PLD.

36 MATERIALS SCIENCE↗

MicroFisher: Fungal taxonomic classification for metatranscriptomic and metagenomic data using multiple short hypervariable markers

AbstractProfiling the taxonomic and functional composition of microbes using metagenomic (MG) and metatranscriptomic (MT) sequencing is advancing our understanding of microbial functions. However, the sensitivity and accuracy of microbial classification using genome– or core protein-based approaches, especially the classification of eukaryotic organisms, is limited by the availability of genomes and the resolution of sequence databases. To address this, we propose the MicroFisher, a novel approach that applies multiple hypervariable marker genes to profile fungal communities from MGs and MTs. This approach utilizes the hypervariable regions of ITS and large subunit (LSU) rRNA genes for fungal identification with high sensitivity and resolution. Simultaneously, we propose a computational pipeline (MicroFisher) to optimize and integrate the results from classifications using multiple hypervariable markers. To test the performance of our method, we applied MicroFisher to the synthetic community profiling and found high performance in fungal prediction and abundance estimation. In addition, we also used MGs from forest soil and MTs of root eukaryotic microbes to test our method and the results showed that MicroFisher provided more accurate profiling of environmental microbiomes compared to other classification tools. Overall, MicroFisher serves as a novel pipeline for classification of fungal communities from MGs and MTs.

Wang, Haihua↗

A Comparative Study of Physics‐Informed and Data‐Driven Neural Networks for Compound Flood Simulation at River‐Ocean Interfaces: A Case Study of Hurricane Irene

Simulating compound flooding (CF) at the river-ocean interface within large-scale Earth System Models (ESMs) presents significant challenges due to complex interactions between river discharge, storm surge, and tides. This study assesses the comparative advantages of physics-informed and data-driven machine learning (ML) approaches for enhancing local ESM performance. We systematically compare data-driven neural network models (i.e., CNNs, U-Net, Long Short-Term Memory (LSTM), Gated Recurrent Unit), and physics-informed neural network (PINN) models, including vanilla PINN and a finite-difference-based PINN (FD-PINN). Specifically, FD-PINN is introduced to enhance computational efficiency, accelerating vanilla PINNs by ∼6.5 times while improving accuracy. To enhance data-driven model training, a new data-generation approach is developed to sample historical fluvial and coastal flood events, which ensures a robust data set for extreme event prediction. The models are evaluated using a realistic one-dimensional river domain extracted from an ESM's river mesh and the Hurricane Irene event as an independent test case. Results show that FD-PINN achieves accurate predictions with significantly reduced computational costs relative to vanilla PINNs. Among data-driven models, the best overall performance is achieved by a CNN-LSTM hybrid, which balances accuracy and efficiency. While a fully connected CNN (CNN-FC) provides the best accuracy, it incurs high computational cost. Architectures lacking strong temporal modeling tend to underperform on unseen events. These findings highlight the importance of sequence-aware designs for robust generalization. This study reveals the trade-offs between physics-informed and data-driven models and proposes an adaptive hybrid framework for integrating ML into ESMs to enhance local flood simulations.

Earth Systems Modeling↗

The final WaZP galaxy cluster catalog of the Dark Energy Survey and comparison with SZE data

In this work, we present and characterize the galaxy cluster catalog detected by the WaZP cluster finder, which is not based on red-sequence identification, on the full six years of observations of the Dark Energy Survey (DES-Y6). The full catalog contains over 400k detected clusters with richnesses, Ngals, above 5 and that reach redshifts up to 1.3. We also provide a version of the catalog where the observation depth and richness computation are homogenized to be used for cosmology, containing 33k rich (Ngals >25) clusters. We compare our results with the previous WaZP catalog obtained from the DES first-year data release (DES-Y1). We find that essentially all clusters within the common footprint and depth limit are recovered. The deeper observations on DES-Y6 and the more complete available spectroscopic redshift sample lead to improvements in the redshifts of the clusters, resulting in an average scatter of 1.4% and offset of 0.2%. The optical clusters are also cross-matched with Sunyaev Zel'dovich Effect (SZE) cluster samples detected by the South Pole Telescope (SPT) and the Atacama Cosmology Telescope (ACT). We find that essentially all SZE clusters with reasonable overlapping footprint have a corresponding WaZP cluster. Conversely, 90% of the optical detections with richness greater than 150 have a counterpart in the deeper regions of the SZE surveys. Based on cross-match with the SZE catalogs, we also find that 15-20% of the SZE matched systems have more than one possible WaZP counterpart at the same redshift and within the SZE R500c, indicating possible interacting or unrelaxed systems. Finally, given the optical and SZE beams, WaZP and SZE centerings are found to be consistent. A more detailed study of the SZE-WaZP mass-richness relation will be presented in a separate paper.

Benoist, C. [OCA, Nice, Lab. Lagrange; LIneA, Rio ↗