Search NASA⌕ Search

SEARCH · Search NASA

Results for “reference transcript dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A high-resolution single-molecule sequencing-based Arabidopsis transcriptome using novel methods of Iso-seq analysis

Accurate and comprehensive annotation of transcript sequences is essential for transcript quantification and differential gene and transcript expression analysis. Single-molecule long-read sequencing technologies provide improved integrity of transcript structures including alternative splicing, and transcription start and polyadenylation sites. However, accuracy is significantly affected by sequencing errors, mRNA degradation, or incomplete cDNA synthesis. We present a new and comprehensive Arabidopsis thaliana Reference Transcript Dataset 3 (AtRTD3). AtRTD3 contains over 169,000 transcripts—twice that of the best current Arabidopsis transcriptome and including over 1500 novel genes. Seventy-eight percent of transcripts are from Iso-seq with accurately defined splice junctions and transcription start and end sites. We develop novel methods to determine splice junctions and transcription start and end sites accurately. Mismatch profiles around splice junctions provide a powerful feature to distinguish correct splice junctions and remove false splice junctions. Stratified approaches identify high-confidence transcription start and end sites and remove fragmentary transcripts due to degradation. AtRTD3 is a major improvement over existing transcriptomes as demonstrated by analysis of an Arabidopsis cold response RNA-seq time-series. AtRTD3 provides higher resolution of transcript expression profiling and identifies cold-induced differential transcription start and polyadenylation site usage. AtRTD3 is the most comprehensive Arabidopsis transcriptome currently. It improves the precision of differential gene and transcript expression, differential alternative splicing, and transcription start/end site usage analysis from RNA-seq data. The novel methods for identifying accurate splice junctions and transcription start/end sites are widely applicable and will improve single-molecule sequencing analysis from any species.

transcription start and end sites↗

A consensus-based ensemble approach to improve transcriptome assembly

Systems-level analyses, such as differential gene expression analysis, co-expression analysis, and metabolic pathway reconstruction, depend on the accuracy of the transcriptome. Multiple tools exist to perform transcriptome assembly from RNAseq data. However, assembling high quality transcriptomes is still not a trivial problem. This is especially the case for non-model organisms where adequate reference genomes are often not available. Different methods produce different transcriptome models and there is no easy way to determine which are more accurate. Furthermore, having alternative-splicing events exacerbates such difficult assembly problems. While benchmarking transcriptome assemblies is critical, this is also not trivial due to the general lack of true reference transcriptomes. In this study, we first provide a pipeline to generate a set of the simulated benchmark transcriptome and corresponding RNAseq data. Using the simulated benchmarking datasets, we compared the performance of various transcriptome assembly approaches including both de novo and genome-guided methods. The results showed that the assembly performance deteriorates significantly when alternative transcripts (isoforms) exist or for genome-guided methods when the reference is not available from the same genome. To improve the transcriptome assembly performance, leveraging the overlapping predictions between different assemblies, we present a new consensus-based ensemble transcriptome assembly approach, ConSemble. Without using a reference genome, ConSemble using four de novo assemblers achieved an accuracy up to twice as high as any de novo assemblers we compared. When a reference genome is available, ConSemble using four genome-guided assemblies removed many incorrectly assembled contigs with minimal impact on correctly assembled contigs, achieving higher precision and accuracy than individual genome-guided methods. Furthermore, ConSemble using de novo assemblers matched or exceeded the best performing genome-guided assemblers even when the transcriptomes included isoforms. We thus demonstrated that the ConSemble consensus strategy both for de novo and genome-guided assemblers can improve transcriptome assembly. The RNAseq simulation pipeline, the benchmark transcriptome datasets, and the script to perform the ConSemble assembly are all freely available from: http://bioinfolab.unl.edu/emlab/consemble/.

59 BASIC BIOLOGICAL SCIENCES↗

Transcribing Air Traffic Control System Command Center Planning Telecons Using Cloud-Based Automatic Speech Recognition

This paper addresses the challenge of using Automatic Speech Recognition (ASR) technology to transcribe regular teleconferences that happen between FAA Air Traffic Control System Command Center (ATCSCC) planners, stakeholders and air users. These planning teleconferences (aka telecons or planning webinars) are an integral part of managing air traffic in the U.S. National Airspace System (NAS). In particular, the meetings facilitate the creation and modification of various traffic management initiatives (TMIs), that are used to regulate the flow of air traffic. This is typically a human intensive process, requiring specialists to listen to the entire meeting audio (10-20 minutes duration) and inferring the state of the NAS (e.g., weather phenomenon) that was discussed. It would be advantageous to have digital transcripts of the audio and have useful information (e.g., related to TMIs) automatically extracted from the transcripts. In this regard, we are exploring the adoption of state-of-the-art speech to text and Natural Language Processing (NLP) tools that will achieve our objective of digitizing the webinar audio. Unfortunately, the highly technical phraseology present in the audio and limited data availability for model building make ASR difficult. To overcome this challenge, we have taken the critical first step in creating a human transcription dataset from ~20 hours of speech in the ATCSCC audio with the help of subject matter experts. A novelty of our work is the creation of a ground truth transcription dataset for ATCSCC teleconference webinars, which is particularly important for Aviation domain-specific NLP tasks. Using Microsoft Speech Studio, a cloud-based ASR platform, we have fine-tuned the English pre-trained ASR models (available in speech studio) and achieved an average word error rate (WER) of 6.81%. The baseline ASR also provides a digital version of each planning webinar, making it accessible and text-searchable for future references. Additionally, the transcriptions can serve as a bridge between raw audio data and a range of text-based NLP tasks, such as named entity recognition (NER) and intent classification, potentially enhancing the digital footprint of the webinars and other connected data sources. Our work has several potential applications. Firstly, the transcriptions can be analyzed to understand the complex decision process of creating, implementing and modifying TMIs and may also contribute to TMI prediction services. Secondly, our dataset and model can be used to develop more accurate ASR systems for aviation-specific language, which can bring about digital communication in the aviation industry (and aid current “voice only” communications, which are inherently error-prone). Lastly, the transcriptions themselves can be used as a valuable resource for training other NLP models.

Stephen S. B. Clarke↗

Structural models and functional annotations for the Sphagnum divinum proteome

This dataset contains the structural models for the primary transcripts of the Sphagnum divinum proteome. Additionally, for a subset of these proteins, sequence and structural alignment results are provided. This dataset represents the most thorough structural study of a Sphagnum species, also known as peat mosses, by providing three-dimensional atomic resolution structures of the majority of the encoded proteins as well as structural alignment results used in the application of annotating the proteome. References (DOI) AlphaFold v2 Monomer: https://doi.org/10.1038/s41586-021-03819-2. References (DOI) US-align2: https://doi.org/10.1038/s41592-022-01585-1

59 BASIC BIOLOGICAL SCIENCES↗

Temporal regulation of cold transcriptional response in switchgrass

Switchgrass low-land ecotypes have significantly higher biomass but lower cold tolerance compared to up-land ecotypes. Understanding the molecular mechanisms underlying cold response, including the ones at transcriptional level, can contribute to improving tolerance of high-yield switchgrass under chilling and freezing environmental conditions. Here, by analyzing an existing switchgrass transcriptome dataset, the temporal cis- regulatory basis of switchgrass transcriptional response to cold is dissected computationally. We found that the number of cold-responsive genes and enriched Gene Ontology terms increased as duration of cold treatment increased from 30 min to 24 hours, suggesting an amplified response/cascading effect in cold-responsive gene expression. To identify genomic sequences likely important for regulating cold response, machine learning models predictive of cold response were established using k -mer sequences enriched in the genic and flanking regions of cold-responsive genes but not non-responsive genes. These k -mers, referred to as putative cis -regulatory elements (pCREs) are likely regulatory sequences of cold response in switchgrass. There are in total 655 pCREs where 54 are important in all cold treatment time points. Consistent with this, eight of 35 known cold-responsive CREs were similar to top-ranked pCREs in the models and only these eight were important for predicting temporal cold response. More importantly, most of the top-ranked pCREs were novel sequences in cold regulation. Our findings suggest additional sequence elements important for cold-responsive regulation previously not known that warrant further studies.

60 APPLIED LIFE SCIENCES↗

Generating Co-expression Networks for Three Cyanobacteria: Synechococcus sp. PCC 7942, Synechococcus sp. PCC 7002, Synechocystis sp. PCC 6803

Cyanobacteria are photosynthetic organisms capable of high growth rate and represent a promising bioplatform for harnessing the sun’s energy to make biofuel. Additionally, the process of photosynthesis absorbs CO2 from the environment. Understanding the metabolic processes involved in photosynthesis could lead to solutions to the recent rise of CO2 concentration in Earth’s atmosphere and the associated climate change. More research on the transcriptional regulation of these cells is needed to learn how to harness the untapped potential of cyanobacteria for these applications. Transcriptional analysis via RNA-seq provides an understanding of how gene expression changes at the mRNA level under diverse growing conditions. I systematically collected and analyzed RNA-Seq data obtained under a variety of conditions and available on the NCBI database for three cyanobacteria model organisms: Synechococcus elongatus sp. PCC 7942, Synechococcus sp. PCC 7002, and Synechocystis sp. PCC 6803. For each organism, the data was mapped to a reference genome to characterize the RNA expression profile. Samples were checked for quality based on the number of reads and the correlation of the expression profile between labeled replicates. All samples were transformed into transcripts per million reads, followed by a log transformation to account for the wide range of sample sizes. Gene co-expression networks were generated and analyzed for each species using Cytoscape. These networks provide a base level of gene expression for each species. The network topology and high-betweeness nodes of these networks need to be analyzed further to provide insight on potential ways to harness cyanobacteria genetics. Additionally, these datasets can be used together to form a core genome network analysis- one that includes only the genes that are homologous between the three species. This project has prepared the way for a more in-depth study on photosynthetic microbes on a genetic level.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

A Pipeline for Assessing the Quality of Rna-Seq Datasets in GeneLab

Transcriptome profiling by RNA sequencing (RNA-seq) is a powerful approach to identify gene expression changes in organisms exposed to unique environments such as spaceflight. One of the challenges of evaluating RNA-seq data both within and across different space-relevant studies is the ability to control for technical differences, including the use of different library preparation kits, sequencing platforms, RNA yield, and person-to-person variation. To help address this issue, the National Institute of Standards and Technology (NIST, nist.gov) initiated a consortium, at the request of industry and academia, to develop a set of controls for gene expression measurements. The result was a set of 92 unlabeled, polyadenylated transcripts that range from 250 – 2,000 nucleotides in length to mimic natural eukaryotic mRNAs. These External RNA Controls Consortium (ERCC) genes can be used in any RNA-seq experiment, by adding known concentrations of the ERCC genes to samples after RNA extraction, to offer a standard measurement for data comparison. At NASA GeneLab, we employ these controls as part of our standard operating procedures for every in-house RNA-seq study to assess the limit of detection, dynamic range, and power of differential expression analysis both within and across experiments. Here we will discuss the use, benefits, and limitations of ERCC genes and other types of controls, such as universal RNA references, to generate quality control information for RNA-seq studies conducted at GeneLab.

GeneLab↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Snowex 2017 Community Snow Depth Measurements: A Quality-Controlled, Georeferenced Product

Snow depth was one of the core ground measurements required to validate remotely-sensed data collected during SnowEx Year 1, which occurred in Colorado. The use of a single, common protocol was fundamental to produce a community reference dataset of high quality. Most of the nearly 100 Grand Mesa and Senator Beck Basin SnowEx ground crew participants contributed to this crucial dataset during 6-25 February 2017. Snow depths were measured along ~300 m transects, whose locations were determined according to a random-stratified approach using snowfall and tree-density gradients. Two-person teams used snowmobiles, skis, or snowshoes to travel to staked transect locations and to conduct measurements. Depths were measured with a 1-cm incremented probe every 3 meters along transects. In shallow areas of Grand Mesa, depth measurements were also collected with GPS snow-depth probes (a.k.a. MagnaProbes) at ~1-m intervals. During summer 2017, all reference stake positions were surveyed with <10 cm accuracy to improve overall snow depth location accuracy. During the campaign, 193 transects were measured over three weeks at Grand Mesa and 40 were collected over two weeks in Senator Beck Basin, representing more than 27,000 depth values. Each day of the campaign depth measurements were written in waterproof field books and photographed by National Snow and Ice Data Center (NSIDC) participants. The data were later transcribed and prepared for extensive quality assessment and control. Common issues such as protocol errors (e.g., survey in reverse direction), notebook image issues (e.g., halo in the center of digitized picture), and data-entry errors (sloppy writing and transcription errors) were identified and fixed on a point-by-point basis. In addition, we strove to produce a georeferenced product of fine quality, so we calculated and interpolated coordinates for every depth measurement based on surveyed stakes and the number of measurements made per transect. The product has been submitted to NSIDC in csv format. To educate data users, we present the study design and processing steps that have improved the quality and usability of this product. Also, we will address measurement and design uncertainties, which are different in open vs. forest areas.

Brucker, L.↗

Bringing Planetary Science Mission Outreach to the Deaf and Blind Communities

Introduction: Technology for enhancing outreach, like 3D printing, and science communication products, such as videos and podcasts, can be utilized within the planetary science community, especially for the engagement and excitement of current or upcoming planetary exploration missions. However, these communication products can also be further enhanced for the benefit of the blind and deaf communities. While such products may already be readily available, such projects are not easily accessible to blind and/or deaf certified educators, which often rely on making their own resources or do not have the funds to provide such resources (e.g., cost of 3D printers or cost of braille books). The planetary science community can have better practices to reach these broader audiences. Best practices can include transcripts from podcasts, transcripts in videos, and large-font captions. Images on websites and social media accounts should also include alt-text descriptive captions. 3D printing can also enhance planetary science for the blind community, through tactile posters, maps, and pamphlets. Planetary data can also be augmented by providing different tactile geological maps (e.g., topography or various datasets), and audio-visual videos freely available for educators. Visual Engagement: Visual engagement consists of several avenues to consider, the three main themes includes: 1) swag; 2) videos; 3) interactive exploration. Swag can include the fun visual take-home materials, such as stickers, posters, bookmarks, etc. Videos can include educational-specific videos (available freely via YouTube or by other educational-specific streaming avenues, such as Nebula or Curiosity Stream), provided they have Closed Captioning (CC). The use of QR codes to such videos or websites can also benefit to being added on swag. Interactive exploration can also be sub-divided by different types of engagement. A popular and still fairly new technology for public engagement is the use of virtual reality (VR). While this has been mainly for martian and lunar surface exploration [1], the deaf communities can benefit from VR through a more extensive look at our solar system and beyond (for example, a VR experience of the flight path, or visual map of the heliosphere/dynamics of our Sun). Audio Engagement: Audio tools can also be a useful avenue of communication, especially for the blind communities. Audio archiving can certainly be transcripts from the video engagements, but also the use of podcasts can also be a benefit. Podcasting can take on two forms: 1) interview engagement; and 2) update engagement. For interviews, scientists can communicate with STEM-specific podcast platforms to make other listener-bases aware of what is going on with a specific mission. For update-type communication, missions may opt to have an archived podcast of news, updates, and the teams involved. The most important aspect of such podcasts would be for the need of complimentary transcripts (including descriptive transcripts if sounds are included), and the limited use of jargon. Research Engagement: There have been several examples of involving the blind and low-vision communities in citizen science, such as through the NASA Heliophysics division. Examples include the NASA PUNCH (Polarimeter to Unify the Corona and Heliosphere) mission led by the Southwest Research Institute [2], which include blind and visually-impaired citizens to assist in the Sun’s coronal rhythms.” Another example is the Eclipse Soundscapes: Citizen Science Project (ES:CSP), which documents observations of acoustical changes of nature and ecosystems during solar eclipse events [3]. Inclusivity: A major theme that is necessary for public engagement is inclusivity and the awareness of reaching broader audiences. Outreach to include hearing/seeing impaired communities are still lacking in the sciences. There are several opportunities that the planetary sciences could take. Other projects that have emerged from the space sciences include adding transcripts to visual engagement [4], and the use of 3D printing for the visually-impaired [5]. References: [1] Olgin, J. (2020) 51st LPSC, Abstract 2137. [2] https://scitechdaily.com/outreach-for-nasa-punch-mission-embraces-ancient-and-modern-sun-watching-theme/ [3] https://science.nasa.gov/science-activation-team/eclipse-soundscapes [4] NASA International Observe the Moon Night (Blind and Deaf Accessible), Youtube Video ( https://www.youtube.com/watch?v=neHCfg0S3-Q) [5] Richardson, J., et al. (2018) AGU Fall Meeting, Abstract ED23F-0963.

C J Ahrens↗