Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

Bayesian estimation of HIV acquisition dates for prevention trials

Accurate timing estimates of when participants acquire HIV in HIV prevention trials are necessary for determining antibody levels at acquisition. The Antibody-Mediated Prevention (AMP) Studies showed that a passively administered broadly neutralizing antibody can prevent the acquisition of HIV from a neutralization-sensitive virus. We developed a pipeline for estimating the date of detectable HIV acquisition (DDA) in AMP Study participants using diagnostic and viral sequence data. Using a Bayesian strategy that combines three streams of data (REN [rev/vpu/env/Δnef] sequence, GP [gag/Δpol] sequence, and diagnostic) where their 95% credible intervals overlap based on pre-specified criteria and decision rules. We evaluated the performance of our AMP pipeline using PacBio viral sequence data from 41 participants across two prospective acute HIV acquisition cohort studies, FRESH and RV217, with twice-weekly sampling. These cohort studies enrolled young women in South Africa and men and women in Kenya and Thailand, respectively, with a high likelihood of HIV acquisition. In evaluating performance, “true DDA” was the center of bounds between last-negative and first-positive RNA diagnostic tests (median time 4 days, range 2–7 days); bias was the mean difference between estimated and true DDA. Using diagnostic data alone yielded timing estimates with a bias of 2.4 days and root mean square error (RMSE) of 7.9 days. These results were improved using sequence + diagnostic data (bias 1.5 days, RMSE 6.9 days), as well as by restricting sequence-based estimation to samples from ≤5 weeks post-DDA (bias 0.2 days, RMSE 7.8 days).

59 BASIC BIOLOGICAL SCIENCES

Omega flight-test data reduction sequence

Computer programs for Omega data conversion, summary, and preparation for distribution are presented. Program logic and sample data formats are included, along with operational instructions for each program. Flight data (or data collected in flight format in the laboratory) is provided by the Ohio University Omega receiver base in the form of 6-bit binary words representing the phase of an Omega station with respect to the receiver's local clock. All eight Omega stations are measured in each 10-second Omega time frame. In addition, an event-marker bit and a time-slot D synchronizing bit are recorded. Program FDCON is used to remove data from the flight recorder tape and place it on data-processing cards for later use. Program FDSUM provides for computer plotting of selected LOP's, for single-station phase plots, and for printout of basic signal statistics for each Omega channel. Mean phase and standard deviation are printed, along with data from which a phase distribution can be plotted for each Omega station. Program DACOP simply copies the Omega data deck a controlled number of times, for distribution to users.

Lilley, R. W.

Archaeal phylogeny: reexamination of the phylogenetic position of Archaeoglobus fulgidus in light of certain composition-induced artifacts

A major and too little recognized source of artifact in phylogenetic analysis of molecular sequence data is compositional difference among sequences. The problem becomes particularly acute when alignments contain ribosomal RNAs from both mesophilic and thermophilic species. Among prokaryotes the latter are considerably higher in G + C content than the former, which often results in artificial clustering of thermophilic lineages and their being placed artificially deep in phylogenetic trees. In this communication we review archaeal phylogeny in the light of this consideration, focusing in particular on the phylogenetic position of the sulfate reducing species Archaeoglobus fulgidus, using both 16S rRNA and 23S rRNA sequences. The analysis shows clearly that the previously reported deep branching of the A. fulgidus lineage (very near the base of the euryarchaeal side of the archaeal tree) is incorrect, and that the lineage actually groups with a previously recognized unit that comprises the Methanomicrobiales and extreme halophiles.

NASA Discipline Exobiology

The rate and efficiency of high-mass star formation along the Hubble sequence

Data obtained with IRAS are used to compare and contrast the global star formation rates for a galactic sample which represents essentially all known noninteracting spiral and lenticular galaxies within 40 Mpc. The distribution of 60 micron luminosity is similar for spirals of types Sa-Scd inclusively, although the luminosities of the very early and very late types are, on average, one order of magnitude lower. High-mass star formation rates are similar for early, intermediate, and late type spirals, and the average high-mass star formation rate per unit molecular gas mass is independent of type for spiral galaxies. A remarkable homogeneity exists in the high-mass star-forming capabilities of spiral galaxies, particularly among the Sa-Scd types. The Hubble sequence is therefore not a sequence in the present-day rate or production efficiency of high-mass stars.

Devereux, Nicholas A.

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab

Cascade Error Projection with Low Bit Weight Quantization for High Order Correlation Data

In this paper, we reinvestigate the solution for chaotic time series prediction problem using neural network approach. The nature of this problem is such that the data sequences are never repeated, but they are rather in chaotic region. However, these data sequences are correlated between past, present, and future data in high order. We use Cascade Error Projection (CEP) learning algorithm to capture the high order correlation between past and present data to predict a future data using limited weight quantization constraints. This will help to predict a future information that will provide us better estimation in time for intelligent control system. In our earlier work, it has been shown that CEP can sufficiently learn 5-8 bit parity problem with 4- or more bits, and color segmentation problem with 7- or more bits of weight quantization. In this paper, we demonstrate that chaotic time series can be learned and generalized well with as low as 4-bit weight quantization using round-off and truncation techniques. The results show that generalization feature will suffer less as more bit weight quantization is available and error surfaces with the round-off technique are more symmetric around zero than error surfaces with the truncation technique. This study suggests that CEP is an implementable learning technique for hardware consideration.

Duong, Tuan A.

Delay Tolerant, Radio Frequency Identification (RFID )-enabled Sensing

Radio Frequency Identification (RFID) technology offers a completely passive method to transmit fixed data sequences from an RFID tag, which typically doesn’t have its own power supply, to an interrogator. Radio frequency (RF) energy harvested from the interrogator is rectified by the tag and used to charge an integrated circuit (IC). The IC then modulates the received signal with the data stored on the tag and reflects the energy back to the interrogator. RFID has seen great proliferation in terrestrial inventory management applications, and it has recently made the jump to spaceflight applications onboard the International Space Station, augmenting an existing optical bar‐code infrastructure for tracking supplies. A number of advanced automated logistics management (ALM) concepts employing RFID are currently being developed and evaluated, including so‐called “smart” shelves, cubbies, and trash receptacles using low‐power, embeddable RFID interrogators. An infrastructure where both crew members and robotic assistants, such as autonomous free flyers, are similarly equipped with small RFID interrogators seems likely. It therefore behooves us to consider extending this infrastructure beyond ALM to applications such as low power, embedded sensing. Typically, data on an RFID tag can only be written by an RFID interrogator, using interrogator energy. In recent years, however, a few efforts have focused on using that energy to drive data acquisition from the tag IC, allowing the tag to modify its stored data sequence with sensor data before replying to an interrogator. In this way, the tag can act as a completely passive sensing device. One problem exists with this approach, however: the tag cannot gather data when an interrogator is not present. Thus, strictly passive RFID sensing tags cannot gather data at regular intervals, in the manner of a typical wireless sensor network, without careful, and impractical, planning of mobile interrogator movements. To address this shortcoming, we look to a recent advance in RFID technology which allows an external microprocessor to power the tag IC and write directly into its RFID memory using a wired serial interface. In this paradigm, data gathering is driven by a small, on‐board power supply (using batteries or harvested energy), and data transfer is provided passively through the RFID interrogation service. Since communication typically consumes the lion’s share of power in WSNs, such a technique has the potential to enable extremely long‐lived, embedded wireless sensing when used with extremely low‐current microcontrollers. Since the communication channel is only open when an interrogator is present and actively interrogating the RFID sensing tag, transport of periodically‐sampled sensor data presents itself as a delay/disruption‐tolerant networking (DTN) problem. In this paper, we present the design of a DTN‐like overlay on the common EPC Global, Class 1, Generation 2 RFID standard. This overlay allows seamless, guaranteed data transfer using the EPC Global protocol, supporting extremely low‐power, embedded sensing using an infrastructure likely to be already in place for ALM applications. We evaluate a prototype implementation of a complete end‐to‐end system using a robotic RFID interrogation agent, and we present future directions for the development of this sensing technique.

Raymond S. Wagner

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

Non-NASA Center

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES

Parallel VLSI Architecture

Fermat number transformation convolutes two digital data sequences. Very-large-scale integration (VLSI) applications, such as image and radar signal processing, X-ray reconstruction, and spectrum shaping, linear convolution of two digital data sequences of arbitrary lenghts accomplished using Fermat number transform (ENT).

Truong, T. K.

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R 2 of 0.604 [0.593–0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156–5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

COVID19

Phylogenetic diversity and position of the genus Campylobacter

RNA sequence analysis has been used to examine the phylogenetic position and structure of the genus Campylobacter. A complete 5S rRNA sequence was determined for two strains of Campylobacter jejuni and extensive partial sequences of the 16S rRNA were obtained for several strains of C. jejuni and Wolinella succinogenes. In addition limited partial sequence data were obtained from the 16S rRNAs of isolates of C. coli, C. laridis, C. fetus, C. fecalis, and C. pyloridis. It was found that W. succinogenes is specifically related to, but not included, in the genus Campylobacter as presently constituted. Within the genus significant diversity was noted. C. jejuni, C. coli and C. laridis are very closely related but the other species are distinctly different from one another. C. pyloridis is without question the most divergent of the Campylobacter isolates examined here and is sufficiently distinct to warrant inclusion in a separate genus. In terms of overall position in bacterial phylogeny, the Campylobacter/Wolinella cluster represents a deep branching most probably located within an expanded version of the Division containing the purple photosynthetic bacteria and their relatives. The Campylobacter/Wolinella cluster is not specifically includable in either the alpha, beta or gamma subdivisions of the purple bacteria.

NASA Discipline Exobiology

The imaging experiment on Pioneer 10

In the period plus or minus 4 days from Jovian pericenter, Pioneer 10 returned data from 153 imaging sequences. Data were obtained with the imaging photopolarimeter, a 2.54-cm-diam and 8.6-cm-focal-length steerable telescope configured as a narrow beam (0.5-mrad) photometer in a spin scan mode of operation. Images were obtained in two spectral bands (390 to 500 nm and 595 to 720 nm), and seven of the most interesting pictures are shown and discussed.

Swindell, W.

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES

Evolution of thermotolerance in hot spring cyanobacteria of the genus Synechococcus

The extension of ecological tolerance limits may be an important mechanism by which microorganisms adapt to novel environments, but it may come at the evolutionary cost of reduced performance under ancestral conditions. We combined a comparative physiological approach with phylogenetic analyses to study the evolution of thermotolerance in hot spring cyanobacteria of the genus Synechococcus. Among the 20 laboratory clones of Synechococcus isolated from collections made along an Oregon hot spring thermal gradient, four different 16S rRNA gene sequences were identified. Phylogenies constructed by using the sequence data indicated that the clones were polyphyletic but that three of the four sequence groups formed a clade. Differences in thermotolerance were observed for clones with different 16S rRNA gene sequences, and comparison of these physiological differences within a phylogenetic framework provided evidence that more thermotolerant lineages of Synechococcus evolved from less thermotolerant ancestors. The extension of the thermal limit in these bacteria was correlated with a reduction in the breadth of the temperature range for growth, which provides evidence that enhanced thermotolerance has come at the evolutionary cost of increased thermal specialization. This study illustrates the utility of using phylogenetic comparative methods to investigate how evolutionary processes have shaped historical patterns of ecological diversification in microorganisms.

Synechococcus Group/classification/growth & develo