Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Mechanistic modeling of in vitro transcription incorporating effects of magnesium pyrophosphate crystallization

The in vitro transcription (IVT) reaction used in the production of messenger RNA vaccines and therapies remains poorly quantitatively understood. Mechanistic modeling of IVT could inform reaction design, scale-up, and control. In this work, we develop a mechanistic model of IVT to include nucleation and growth of magnesium pyrophosphate crystals and subsequent agglomeration of crystals and DNA. To help generalize this model to different constructs, a novel quantitative description is included for the rate of transcription as a function of target sequence length, DNA concentration, and T7 RNA polymerase concentration. The model explains previously unexplained trends in IVT data and quantitatively predicts the effect of adding the pyrophosphatase enzyme to the reaction system. The model is validated on additional literature data showing an ability to predict transcription rates as a function of RNA sequence length.

59 BASIC BIOLOGICAL SCIENCES↗

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES↗

Adapting CLUTCH methodology to multigroup TSUNAMI-3D for eigenvalue sensitivity calculations

The sensitivity of the eigenvalue to uncertainties in nuclear data and its evaluation are important for nuclear criticality safety. TSUNAMI-3D sequences within the SCALE code system offer several options to the user community for calculating eigenvalue sensitivity coefficients with multigroup (MG) and continuous energy (CE) 3D transport capabilities. TSUNAMI-3D sequences implement the adjoint-based perturbation theory with MG KENO code, the Contributon Linked eigenvalue sensitivity/Uncertainty estimation via Track length importance CHaracterization (CLUTCH) method with CE KENO code, and the Iterated Fission Probability (IFP) method with CE KENO and Shift codes. Each method has benefits and limitations depending on the problem that is run. The work presented here aims to adapt the CLUTCH method, which enables the Contributon method's mesh-free, memory-efficient approach for calculating adjoint-weighted tallies for sensitivity calculations, to the MG TSUNAMI-3D sequence. This application would eliminate the explicit adjoint KENO calculation, as well as the memory-consuming mesh flux moment tallies required by the conventional MG TSUNAMI-3D. Smaller memory footprints in the CLUTCH methodology and relatively shorter runtimes in MG KENO transport can make MG TSUNAMI-3D a viable method for some complex problems. Moreover, this adaptation allows MG sensitivity calculations with Shift, ORNL's next-generation high-performance Monte Carlo transport code, which currently does not offer any sensitivity capabilities with MG particle transport simulations. Initial implementation of the new MG TSUNAMI-3D sequence and its preliminary results with a selected critical benchmark experiment in the Verified, Archived Library of Inputs and Data (VALID) are presented in this study.

KENO↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

Design and Characterization of a Transcriptional Repression Toolkit for Plants

Regulation of gene expression is essential for all life. Tools to manipulate the gene expression level have therefore proven to be very valuable in efforts to engineer biological systems. However, there are few well-characterized genetic parts that reduce gene expression in plants, commonly known as transcriptional repressors. We characterized the repression activity of a library consisting of repression motifs from approximately 25% of the members of the largest known family of repressors. Combining sequence information with our trans-regulatory function data, we next generated a library of synthetic transcriptional repression motifs with function predicted in advance. After characterizing our synthetic library, we demonstrated not only that many of our synthetic constructs were functional as repressors but also that our advance predictions of repression strength were better than random guesses. Finally, we assessed the functionality of known transcriptional repression motifs from a wide range of eukaryotes. Our study represents the largest plant repressor motif library experimentally characterized to date, providing unique opportunities for tuning transcription in plants.

59 BASIC BIOLOGICAL SCIENCES↗

A Methodology to Evaluate the Grid Reliability Impact of Oscillations Induced by Large Loads

The rapid growth of hyperscale AI data centers is bringing renewed attention to the reliability risk that sustained forced oscillations pose to bulk power systems, with cyclic computational workloads emerging as a new forcing source. Unlike the broadband, stochastic disturbances from traditional industrial loads such as arc furnaces, AI training and inference facilities can inject large active power swings concentrated at specific frequencies over extended durations - characteristics that existing grid planning practices do not account for. While the North American Electric Reliability Corporation (NERC) has recognized this gap and called for system-level studies of large load interconnections, no standardized methodology exists to screen, simulate, and quantify these risks at the planning stage. This report presents the Risk Assessment Tool for Large Load-induced Events (RATLLE), a Python-based, publicly available script suite developed at the Pacific Northwest National Laboratory to evaluate bulk power system reliability risks from data center-induced oscillations. RATLLE implements a three-module workflow: a screening module that identifies vulnerable interconnection locations and excitable system modes; a simulation module that models cyclic data center load behavior using a commercial positive sequence simulation platform; and an analysis module that computes risk metrics and generates interactive visualization dashboards. The risk metrics, formulated around simulation observables, map oscillation impacts to a three-stage severity scale spanning latent equipment fatigue through imminent cascading failure. The methodology is demonstrated on two Western Electricity Coordinating Council (WECC) system models: a publicly available 240-bus reduced representation and a detailed 2031 Heavy Winter planning case. Case studies illustrate that even modest 50 MW forced oscillations at resonant frequencies can produce wide-area power swings, N-1 security constraint violations, and cascading generator trips through protection actions - outcomes that would not occur under normal operating conditions without oscillations present. The results underscore the need for standardized oscillation impact assessment in large load interconnection studies and provide a reproducible, extensible framework for utilities to adopt or customize within their existing planning workflows.

Biswas, Shuchismita↗

Human Host Cellular Response to HCoV-229E Infection Transcriptomics (ACS-DP1)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5), immortalized human lung fibroblasts cells (MRC5) (MOI5), and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue. Sample data was acquired using an Illumina HiSeq 2000 sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Near-Real-Time Material Tracking: Combining Vis–NIR Spectroscopy with Flow Sensing for Accurate Nd(III) Quantification

A fiber-optic visible–near-infrared (vis–NIR) absorption spectroscopy and flow sensor system has been developed for near-real-time tracking of Nd mass in the effluent stream from a column in a fume hood. The approach leverages two unique data streams and a partial least-squares regression (PLSR) model trained on vis–NIR absorption spectra of Nd(III) (0–1.5 M) in 1 M HNO 3 . In-line volumetric flow rate and vis–NIR spectra are measured in sequence after a chromatography column. The time stamps from each data stream are then synchronized, which allows integrated volumes to be combined with Nd(III) molarities predicted by a PLSR model to accurately calculate the Nd mass flowing through the column. This integrated measurement provides instantaneous mass flow and accumulates these data over time to obtain the total mass processed. The methodology developed in this study contributes critical technical infrastructure to improve monitoring capabilities to support chemical separations and the production of strategic materials and isotopes.

Irvine, Sawyer B. [Oak Ridge National Laboratory (↗

Fuel Property Effects on Stochastic Preignition Events During Engine Load Transitions

Stochastic preignition (SPI) is an abnormal combustion phenomenon that can cause catastrophic engine damage. There have been several proposed mechanisms of SPI, where a uniform source is still not certain, however, SPI tendencies have been shown to be influenced by engine operating conditions, oil composition, engine age, and fuel chemical and physical properties. Laboratory research and testing for SPI propensity is challenging given the stochastic nature of events, as well as the potential for significant degradation of the engine platform and measuring equipment over time. Thus, SPI specific experiments are generally conducted under either sustained or cyclic patterning of steady-state operating conditions to avoid the influence of transient engine boundary conditions on test parameters of interest (e.g. oil additive package, fuel properties, engine speed/load, etc.). In this work a cyclically varying SPI test sequence involves a 5 min engine warmup period at a low engine load of around 4 bar gross indicated mean effective pressure (IMEPg), followed by a transition to high load (~20 bar IMEPg) at a constant 2000 rev/min engine speed for a total of 25 min. This individual test sequence load schedule is then sequentially repeated 10 times to generate significant statistical data for analysis. This work examines the influence of fuel chemical and physical properties on SPI tendency during the unsteady portion of the 10-cycle sequence (the first 5 min of the high load operation in each sequence of the loading cycle) which has been discarded from previous analyses due to the uncertainty in engine operating and thermal boundary conditions. Results from this analysis suggest an increasing trend in the ratio of SPI events during the unsteady test period relative to the steady test period with increasing fuel Reid Vapor Pressure (RVP), implying differences in uncontrolled ignition source terms, possibly from, fuel wall interactions and retention during the load transition phase of the test.

Splitter, Derek [ORNL] (ORCID:0000000174044047)↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Circularization of 23S rRNA but not 16S rRNA within archaeal ribosomes

Background Processing of archaeal 16S and 23S rRNAs is believed to involve excision of individual rRNAs from polycistronic precursors, circularization of excised rRNAs, and re-linearization before the incorporation into ribosomes. However, all the knowledge is derived from several isolated species, leaving open the possibility that different processes may occur in other archaeal groups. Results Here, we investigate rRNAs from diverse and mostly uncultivated archaea. Sequencing of total cellular RNA from eight phylum-level lineages indicates that archaeal circular 23S rRNA transcript abundances vastly exceed those of linear counterparts, and linear versions are often undetectable. As the majority of rRNAs derive from mature ribosomes, the data suggest that ribosomes contain circular 23S rRNAs. Thus, we directly sequence RNA extracted from isolated ribosomes of a model archaeon, Methanosarcina acetivorans, and confirm that the 23S rRNAs in the ribosomes are circular. Structural modeling places the 5′ and 3′ ends of the linear precursors of archaeal 23S rRNAs in close proximity to form a GNRA tetraloop (in which N is A, C, G, or U and R is A or G), consistent with their existence as circular molecules. We also confirm the existence of circular 16S rRNA intermediates in transcriptomes of most archaea, yet a circular form is not evident in some distinct archaeal groups, suggesting that certain archaea do not circularize 16S rRNA during processing. Conclusions Our findings uncover unexpected variations in the processing required to generate mature rRNAs and the conformation of functional molecules in archaeal ribosomes.

Archaea↗

Parallel measurement of transcriptomes and proteomes from same single cells using nanodroplet splitting

Single-cell multiomics provides comprehensive insights into gene regulatory networks, cellular diversity, and temporal dynamics. Here, we introduce nanoSPLITS (nanodroplet SPlitting for Linked-multimodal Investigations of Trace Samples), an integrated platform that enables global profiling of the transcriptome and proteome from same single cells via RNA sequencing and mass spectrometry-based proteomics, respectively. Benchmarking of nanoSPLITS demonstrates high measurement precision with deep proteomic and transcriptomic profiling of single-cells. We apply nanoSPLITS to cyclin-dependent kinase 1 inhibited cells and found phospho-signaling events could be quantified alongside global protein and mRNA measurements, providing insights into cell cycle regulation. We extend nanoSPLITS to primary cells isolated from human pancreatic islets, introducing an efficient approach for facile identification of unknown cell types and their protein markers by mapping transcriptomic data to existing large-scale single-cell RNA sequencing reference databases. Accordingly, we establish nanoSPLITS as a multiomic technology incorporating global proteomics and anticipate the approach will be critical to furthering our understanding of biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Development of a kinetic-thermodynamic model for lime-stabilization of Na-bentonite

This study presents the first kinetic model to predict the solid and pore solution composition of Na-bentonite clay reacting with slaked lime over a period of 720 days. The model successfully accounts for most experimental data using a single kinetic rate constant. The following sequence of reactions was predicted by the model: initial rapid dissolution of portlandite within the first 7 days, leading to a decrease in pH and dissolved calcium, and concurrent formation of calcium silicate hydrates (C-S-H: jennite), calcium aluminate hydrate (C-A-H: C₄AH₁₃), calcium aluminosilicate hydrates (stratlingite) and hydrotalcite. After 7 days, jennite and stratlingite are predicted to transform into tobermorite-II, contributing to strength development up to 28 days. From 28 to 90 days, continued montmorillonite dissolution is predicted, along with minor formation of ettringite, partial tobermorite-II dissolution, and precipitation of secondary phases such as albite and talc. Experimentally, portlandite dissolution was confirmed by TGA and XRD and found to be complete within 7 days, in agreement with model predictions. However, other predicted solid-phase transformations (e.g., tobermorite-II formation and dissolution, ettringite, albite, and talc formation) could not be conclusively verified through experimental techniques. Aqueous phase measurements confirmed that the pH and Ca trends in solution, and that equilibrium was reached by 90 days.

Chemical kinetics↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Algorithm 1049: The Delaunay Density Diagnostic

Accurate approximation of a real-valued function depends on two aspects of the available data: the density of inputs within the domain of interest and the variation of the outputs over that domain. There are few methods for assessing whether the density of inputs is sufficient to identify the relevant variations in outputs—i.e., the “geometric scale” of the function—despite the fact that sampling density is closely tied to the success or failure of an approximation method. In this article, we introduce a general purpose, computational approach to detecting the geometric scale of real-valued functions over a fixed domain using a deterministic interpolation technique from computational geometry. The algorithm is intended to work on scalar data in moderate dimensions (2–10). Our algorithm is based on the observation that a sequence of piecewise linear interpolants will converge to a continuous function at a quadratic rate (in L 2 norm) if and only if the data are sampled densely enough to distinguish the feature from noise (assuming sufficiently regular sampling). We present numerical experiments demonstrating how our method can identify feature scale, estimate uncertainty in feature scale, and assess the sampling density for fixed (i.e., static) datasets of input–output pairs. Finally, we include analytical results in support of our numerical findings and have released lightweight code that can be adapted for use in a variety of data science settings.

97 MATHEMATICS AND COMPUTING↗

Proteomics Analysis of Human Contaminant Proteins

Complete characterization of unknowns via proteomics remains challenging. There exist regions of mass spectrometry-based proteomics data where empirical measurements are not attributed to peptides, and/or sequenced peptides from mass spectra are not attributed to any source. These uncharacterized regions are known as the “dark” proteome. Many proteomics tools rely on some a priori knowledge of sample composition; few tools allow for investigation of unknowns without relying on composition assumptions. Further, the potential low abundance of minor traces in these uncharacterized regions can make elucidation of the “dark” proteome challenging. Herein, we describe the development and evaluation of approaches to study the “dark” proteome and move towards an untargeted approach for more complete characterization, namely by studying minor human protein traces in non-human samples and combining that approach with non-human source organism identification without relying on assumptions. Human protein markers, in the form of genetically variant peptides, have been extensively examined in a variety of human matrices, including blood, plasma, and hair, but have yet to be investigated in non-human samples, such as cell cultures, as human contaminant traces. Genetically variant peptides are those that are found in proteins carrying single nucleotide polymorphisms. In this work, we aimed to (1) investigate the feasibility of detecting human contaminant genetically variant peptides (GVPs) in a diverse set of non-human organisms using public proteomics data and a computational pipeline, as well as to (2) develop a combined capability for untargeted source organism characterization and GVP detection. To our knowledge, this is the first report of applying these approaches towards a more complete proteomic characterization of unknowns. We successfully demonstrate the feasibility of broad human contaminant GVP detection in proteomics data, develop a better understanding of GVP detectability, characterize the sample-to-sample variability in GVP detection, and identify a core set of GVPs that can potentially be used as markers indicative of the human contaminant traces portion of the “dark” proteome. Further, we developed and evaluated a combined pipeline, MARLOWE-GVP, that enables both untargeted source organism characterization and GVP detection. We show high accuracy of correct source organism characterization and high degree of similarity of human contaminant GVP detection compared to the conventional approach. Success on both these efforts have allowed us to advance our understanding and characterization of the “dark” proteome.

59 BASIC BIOLOGICAL SCIENCES↗