Search NASA⌕ Search

SEARCH · Search NASA

Results for “reference transcript dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Transcribing Air Traffic Control System Command Center Planning Telecons Using Cloud-Based Automatic Speech Recognition

This paper addresses the challenge of using Automatic Speech Recognition (ASR) technology to transcribe regular teleconferences that happen between FAA Air Traffic Control System Command Center (ATCSCC) planners, stakeholders and air users. These planning teleconferences (aka telecons or planning webinars) are an integral part of managing air traffic in the U.S. National Airspace System (NAS). In particular, the meetings facilitate the creation and modification of various traffic management initiatives (TMIs), that are used to regulate the flow of air traffic. This is typically a human intensive process, requiring specialists to listen to the entire meeting audio (10-20 minutes duration) and inferring the state of the NAS (e.g., weather phenomenon) that was discussed. It would be advantageous to have digital transcripts of the audio and have useful information (e.g., related to TMIs) automatically extracted from the transcripts. In this regard, we are exploring the adoption of state-of-the-art speech to text and Natural Language Processing (NLP) tools that will achieve our objective of digitizing the webinar audio. Unfortunately, the highly technical phraseology present in the audio and limited data availability for model building make ASR difficult. To overcome this challenge, we have taken the critical first step in creating a human transcription dataset from ~20 hours of speech in the ATCSCC audio with the help of subject matter experts. A novelty of our work is the creation of a ground truth transcription dataset for ATCSCC teleconference webinars, which is particularly important for Aviation domain-specific NLP tasks. Using Microsoft Speech Studio, a cloud-based ASR platform, we have fine-tuned the English pre-trained ASR models (available in speech studio) and achieved an average word error rate (WER) of 6.81%. The baseline ASR also provides a digital version of each planning webinar, making it accessible and text-searchable for future references. Additionally, the transcriptions can serve as a bridge between raw audio data and a range of text-based NLP tasks, such as named entity recognition (NER) and intent classification, potentially enhancing the digital footprint of the webinars and other connected data sources. Our work has several potential applications. Firstly, the transcriptions can be analyzed to understand the complex decision process of creating, implementing and modifying TMIs and may also contribute to TMI prediction services. Secondly, our dataset and model can be used to develop more accurate ASR systems for aviation-specific language, which can bring about digital communication in the aviation industry (and aid current “voice only” communications, which are inherently error-prone). Lastly, the transcriptions themselves can be used as a valuable resource for training other NLP models.

Stephen S. B. Clarke↗

A Pipeline for Assessing the Quality of Rna-Seq Datasets in GeneLab

Transcriptome profiling by RNA sequencing (RNA-seq) is a powerful approach to identify gene expression changes in organisms exposed to unique environments such as spaceflight. One of the challenges of evaluating RNA-seq data both within and across different space-relevant studies is the ability to control for technical differences, including the use of different library preparation kits, sequencing platforms, RNA yield, and person-to-person variation. To help address this issue, the National Institute of Standards and Technology (NIST, nist.gov) initiated a consortium, at the request of industry and academia, to develop a set of controls for gene expression measurements. The result was a set of 92 unlabeled, polyadenylated transcripts that range from 250 – 2,000 nucleotides in length to mimic natural eukaryotic mRNAs. These External RNA Controls Consortium (ERCC) genes can be used in any RNA-seq experiment, by adding known concentrations of the ERCC genes to samples after RNA extraction, to offer a standard measurement for data comparison. At NASA GeneLab, we employ these controls as part of our standard operating procedures for every in-house RNA-seq study to assess the limit of detection, dynamic range, and power of differential expression analysis both within and across experiments. Here we will discuss the use, benefits, and limitations of ERCC genes and other types of controls, such as universal RNA references, to generate quality control information for RNA-seq studies conducted at GeneLab.

GeneLab↗

Snowex 2017 Community Snow Depth Measurements: A Quality-Controlled, Georeferenced Product

Snow depth was one of the core ground measurements required to validate remotely-sensed data collected during SnowEx Year 1, which occurred in Colorado. The use of a single, common protocol was fundamental to produce a community reference dataset of high quality. Most of the nearly 100 Grand Mesa and Senator Beck Basin SnowEx ground crew participants contributed to this crucial dataset during 6-25 February 2017. Snow depths were measured along ~300 m transects, whose locations were determined according to a random-stratified approach using snowfall and tree-density gradients. Two-person teams used snowmobiles, skis, or snowshoes to travel to staked transect locations and to conduct measurements. Depths were measured with a 1-cm incremented probe every 3 meters along transects. In shallow areas of Grand Mesa, depth measurements were also collected with GPS snow-depth probes (a.k.a. MagnaProbes) at ~1-m intervals. During summer 2017, all reference stake positions were surveyed with <10 cm accuracy to improve overall snow depth location accuracy. During the campaign, 193 transects were measured over three weeks at Grand Mesa and 40 were collected over two weeks in Senator Beck Basin, representing more than 27,000 depth values. Each day of the campaign depth measurements were written in waterproof field books and photographed by National Snow and Ice Data Center (NSIDC) participants. The data were later transcribed and prepared for extensive quality assessment and control. Common issues such as protocol errors (e.g., survey in reverse direction), notebook image issues (e.g., halo in the center of digitized picture), and data-entry errors (sloppy writing and transcription errors) were identified and fixed on a point-by-point basis. In addition, we strove to produce a georeferenced product of fine quality, so we calculated and interpolated coordinates for every depth measurement based on surveyed stakes and the number of measurements made per transect. The product has been submitted to NSIDC in csv format. To educate data users, we present the study design and processing steps that have improved the quality and usability of this product. Also, we will address measurement and design uncertainties, which are different in open vs. forest areas.

Brucker, L.↗

Bringing Planetary Science Mission Outreach to the Deaf and Blind Communities

Introduction: Technology for enhancing outreach, like 3D printing, and science communication products, such as videos and podcasts, can be utilized within the planetary science community, especially for the engagement and excitement of current or upcoming planetary exploration missions. However, these communication products can also be further enhanced for the benefit of the blind and deaf communities. While such products may already be readily available, such projects are not easily accessible to blind and/or deaf certified educators, which often rely on making their own resources or do not have the funds to provide such resources (e.g., cost of 3D printers or cost of braille books). The planetary science community can have better practices to reach these broader audiences. Best practices can include transcripts from podcasts, transcripts in videos, and large-font captions. Images on websites and social media accounts should also include alt-text descriptive captions. 3D printing can also enhance planetary science for the blind community, through tactile posters, maps, and pamphlets. Planetary data can also be augmented by providing different tactile geological maps (e.g., topography or various datasets), and audio-visual videos freely available for educators. Visual Engagement: Visual engagement consists of several avenues to consider, the three main themes includes: 1) swag; 2) videos; 3) interactive exploration. Swag can include the fun visual take-home materials, such as stickers, posters, bookmarks, etc. Videos can include educational-specific videos (available freely via YouTube or by other educational-specific streaming avenues, such as Nebula or Curiosity Stream), provided they have Closed Captioning (CC). The use of QR codes to such videos or websites can also benefit to being added on swag. Interactive exploration can also be sub-divided by different types of engagement. A popular and still fairly new technology for public engagement is the use of virtual reality (VR). While this has been mainly for martian and lunar surface exploration [1], the deaf communities can benefit from VR through a more extensive look at our solar system and beyond (for example, a VR experience of the flight path, or visual map of the heliosphere/dynamics of our Sun). Audio Engagement: Audio tools can also be a useful avenue of communication, especially for the blind communities. Audio archiving can certainly be transcripts from the video engagements, but also the use of podcasts can also be a benefit. Podcasting can take on two forms: 1) interview engagement; and 2) update engagement. For interviews, scientists can communicate with STEM-specific podcast platforms to make other listener-bases aware of what is going on with a specific mission. For update-type communication, missions may opt to have an archived podcast of news, updates, and the teams involved. The most important aspect of such podcasts would be for the need of complimentary transcripts (including descriptive transcripts if sounds are included), and the limited use of jargon. Research Engagement: There have been several examples of involving the blind and low-vision communities in citizen science, such as through the NASA Heliophysics division. Examples include the NASA PUNCH (Polarimeter to Unify the Corona and Heliosphere) mission led by the Southwest Research Institute [2], which include blind and visually-impaired citizens to assist in the Sun’s coronal rhythms.” Another example is the Eclipse Soundscapes: Citizen Science Project (ES:CSP), which documents observations of acoustical changes of nature and ecosystems during solar eclipse events [3]. Inclusivity: A major theme that is necessary for public engagement is inclusivity and the awareness of reaching broader audiences. Outreach to include hearing/seeing impaired communities are still lacking in the sciences. There are several opportunities that the planetary sciences could take. Other projects that have emerged from the space sciences include adding transcripts to visual engagement [4], and the use of 3D printing for the visually-impaired [5]. References: [1] Olgin, J. (2020) 51st LPSC, Abstract 2137. [2] https://scitechdaily.com/outreach-for-nasa-punch-mission-embraces-ancient-and-modern-sun-watching-theme/ [3] https://science.nasa.gov/science-activation-team/eclipse-soundscapes [4] NASA International Observe the Moon Night (Blind and Deaf Accessible), Youtube Video ( https://www.youtube.com/watch?v=neHCfg0S3-Q) [5] Richardson, J., et al. (2018) AGU Fall Meeting, Abstract ED23F-0963.

C J Ahrens↗