Search NASA⌕ Search

SEARCH · Search NASA

Results for “Speech Intelligibility”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Overview of Artificial Intelligence (AI) at NASA Goddard

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few. During the Tour, we will show a few examples of the current AI activities at NASA Goddard.

Le Moigne, Jacqueline↗

Multichannel spatial auditory display for speech communications

A spatial auditory display for multiple speech communications was developed at NASA/Ames Research Center. Input is spatialized by the use of simplified head-related transfer functions, adapted for FIR filtering on Motorola 56001 digital signal processors. Hardware and firmware design implementations are overviewed for the initial prototype developed for NASA-Kennedy Space Center. An adaptive staircase method was used to determine intelligibility levels of four-letter call signs used by launch personnel at NASA against diotic speech babble. Spatial positions at 30 degrees azimuth increments were evaluated. The results from eight subjects showed a maximum intelligibility improvement of about 6-7 dB when the signal was spatialized to 60 or 90 degrees azimuth positions.

NASA Discipline Space Human Factors↗

Automatic speech recognition predicts contemporaneous earthquake fault displacement

Abstract Significant progress has been made in probing the state of an earthquake fault by applying machine learning to continuous seismic waveforms. The breakthroughs were originally obtained from laboratory shear experiments and numerical simulations of fault shear, then successfully extended to slow-slipping faults. Here we apply the Wav2Vec-2.0 self-supervised framework for automatic speech recognition to continuous seismic signals emanating from a sequence of moderate magnitude earthquakes during the 2018 caldera collapse at the Kīlauea volcano on the island of Hawai’i. We pre-train the Wav2Vec-2.0 model using caldera seismic waveforms and augment the model architecture to predict contemporaneous surface displacement during the caldera collapse sequence, a proxy for fault displacement. We find the model displacement predictions to be excellent. The model is adapted for near-future prediction information and found hints of prediction capability, but the results are not robust. The results demonstrate that earthquake faults emit seismic signatures in a similar manner to laboratory and numerical simulation faults, and artificial intelligence models developed for encoding audio of speech may have important applications in studying active fault zones.

58 GEOSCIENCES↗

Intelligent interfaces for expert systems

Vital to the success of an expert system is an interface to the user which performs intelligently. A generic intelligent interface is being developed for expert systems. This intelligent interface was developed around the in-house developed Expert System for the Flight Analysis System (ESFAS). The Flight Analysis System (FAS) is comprised of 84 configuration controlled FORTRAN subroutines that are used in the preflight analysis of the space shuttle. In order to use FAS proficiently, a person must be knowledgeable in the areas of flight mechanics, the procedures involved in deploying a certain payload, and an overall understanding of the FAS. ESFAS, still in its developmental stage, is taking into account much of this knowledge. The generic intelligent interface involves the integration of a speech recognizer and synthesizer, a preparser, and a natural language parser to ESFAS. The speech recognizer being used is capable of recognizing 1000 words of connected speech. The natural language parser is a commercial software package which uses caseframe instantiation in processing the streams of words from the speech recognizer or the keyboard. The systems configuration is described along with capabilities and drawbacks.

Villarreal, James A.↗

Experimental Test-Bed for Intelligent Passive Array Research

This document describes the test-bed designed for the investigation of passive direction finding, recognition, and classification of speech and sound sources using sensor arrays. The test-bed forms the experimental basis of the Intelligent Small-Scale Spatial Direction Finder (ISS-SDF) project, aimed at furthering digital signal processing and intelligent sensor capabilities of sensor array technology in applications such as rocket engine diagnostics, sensor health prognostics, and structural anomaly detection. This form of intelligent sensor technology has potential for significant impact on NASA exploration, earth science and propulsion test capabilities. The test-bed consists of microphone arrays, power and signal distribution modules, web-based data acquisition, wireless Ethernet, modeling, simulation and visualization software tools. The Acoustic Sensor Array Modeler I (ASAM I) is used for studying steering capabilities of acoustic arrays and testing DSP techniques. Spatial sound distribution visualization is modeled using the Acoustic Sphere Analysis and Visualization (ASAV-I) tool.

Solano, Wanda M.↗

An Intelligent Computer-aided Training System (CAT) for Diagnosing Adult Illiterates: Integrating NASA Technology into Workplace Literacy

An important part of NASA's mission involves the secondary application of its technologies in the public and private sectors. One current application being developed is The Adult Literacy Evaluator, a simulation-based diagnostic tool designed to assess the operant literacy abilities of adults having difficulties in learning to read and write. Using Intelligent Computer-Aided Training (ICAT) system technology in addition to speech recognition, closed-captioned television (CCTV), live video and other state-of-the-art graphics and storage capabilities, this project attempts to overcome the negative effects of adult literacy assessment by allowing the client to interact with an intelligent computer system which simulates real-life literacy activities and materials and which measures literacy performance in the actual context of its use. The specific objectives of the project are as follows: (1) to develop a simulation-based diagnostic tool to assess adults' prior knowledge about reading and writing processes in actual contexts of application; (2) to provide a profile of readers' strengths and weaknesses; and (3) to suggest instructional strategies and materials which can be used as a beginning point for remediation. In the first and development phase of the project, descriptions of literacy events and environments are being written and functional literacy documents analyzed for their components. From these descriptions, scripts are being generated which define the interaction between the student, an on-screen guide and the simulated literacy environment.

Yaden, David B., Jr.↗

A variable rate speech compressor for mobile applications

One of the most promising speech coder at the bit rate of 9.6 to 4.8 kbits/s is CELP. Code Excited Linear Prediction (CELP) has been dominating 9.6 to 4.8 kbits/s region during the past 3 to 4 years. Its set back however, is its expensive implementation. As an alternative to CELP, the Base-Band CELP (CELP-BB) was developed which produced good quality speech comparable to CELP and a single chip implementable complexity as reported previously. Its robustness was also improved to tolerate errors up to 1.0 pct. and maintain intelligibility up to 5.0 pct. and more. Although, CELP-BB produces good quality speech at around 4.8 kbits/s, it has a fundamental problem when updating the pitch filter memory. A sub-optimal solution is proposed for this problem. Below 4.8 kbits/s, however, CELP-BB suffers from noticeable quantization noise as a result of the large vector dimensions used. Efficient representation of speech below 4.8 kbits/s is reported by introducing Sinusoidal Transform Coding (STC) to represent the LPC excitation which is called Sine Wave Excited LPC (SWELP). In this case, natural sounding good quality synthetic speech is obtained at around 2.4 kbits/s.

Yeldener, S.↗

AI foundation models for experimental fusion tasks

Artificial Intelligence (AI) foundation models, while successful in various domains of language, speech, and vision, have not been adopted in production for fusion energy experiments. This brief paper presents how AI foundation models can be used for fusion energy diagnostics, enabling, for example, visual automated logbooks to provide greater insights into chains of plasma events in a discharge, in time for between-shot analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

An intelligent multi-media human-computer dialogue system

Sophisticated computer systems are being developed to assist in the human decision-making process for very complex tasks performed under stressful conditions. The human-computer interface is a critical factor in these systems. The human-computer interface should be simple and natural to use, require a minimal learning period, assist the user in accomplishing his task(s) with a minimum of distraction, present output in a form that best conveys information to the user, and reduce cognitive load for the user. In pursuit of this ideal, the Intelligent Multi-Media Interfaces project is devoted to the development of interface technology that integrates speech, natural language text, graphics, and pointing gestures for human-computer dialogues. The objective of the project is to develop interface technology that uses the media/modalities intelligently in a flexible, context-sensitive, and highly integrated manner modelled after the manner in which humans converse in simultaneous coordinated multiple modalities. As part of the project, a knowledge-based interface system, called CUBRICON (CUBRC Intelligent CONversationalist) is being developed as a research prototype. The application domain being used to drive the research is that of military tactical air control.

Neal, J. G.↗

Multi-channel spatial auditory display for speech communications

A spatial auditory display for multiple speech communications was developed at NASA-Ames Research Center. Input is spatialized by use of simplified head-related transfer functions, adapted for FIR filtering on Motorola 56001 digital signal processors. Hardware and firmware design implementations are overviewed for the initial prototype developed for NASA-Kennedy Space Center. An adaptive staircase method was used to determine intelligibility levels of four letter call signs used by launch personnel at NASA, against diotic speech babble. Spatial positions at 30 deg azimuth increments were evaluated. The results from eight subjects showed a maximal intelligibility improvement of about 6 to 7 dB when the signal was spatialized to 60 deg or 90 deg azimuth positions.

Begault, Durand↗

Multichannel Spatial Auditory Display for Speed Communications

A spatial auditory display for multiple speech communications was developed at NASA/Ames Research Center. Input is spatialized by the use of simplifiedhead-related transfer functions, adapted for FIR filtering on Motorola 56001 digital signal processors. Hardware and firmware design implementations are overviewed for the initial prototype developed for NASA-Kennedy Space Center. An adaptive staircase method was used to determine intelligibility levels of four-letter call signs used by launch personnel at NASA against diotic speech babble. Spatial positions at 30 degree azimuth increments were evaluated. The results from eight subjects showed a maximum intelligibility improvement of about 6-7 dB when the signal was spatialized to 60 or 90 degree azimuth positions.

Begault, Durand R.↗

Voice intelligibility in satellite mobile communications

An amplitude control technique is reported that equalizes low level phonemes in a satellite narrow band FM voice communication system over channels having low carrier to noise ratios. This method presents at the transmitter equal amplitude phonemes so that the low level phonemes, when they are transmitted over the noisey channel, are above the noise and contribute to output intelligibility. The amplitude control technique provides also for squelching of noise when speech is not being transmitted.

Wishna, S.↗

Exploring Multimodal Interactions in Human-Autonomy Teaming Using a Natural User Interface

The creation of a multimodal, natural user interface to facilitate multi-agent interaction is essential to establishing trust among human and machine teammates in multi-agent systems. Trust is being researched, along with trustworthiness, as a path to certification of autonomous systems by the Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR) project at NASA. The Autonomous Mission Experimental Logistics Interactive Assistant (AMELIA) is a natural user interface that enables multimodal interaction and is designed for rapid mission planning. AMELIA is an intelligent system that considers the user’s preferred communication strategies, as well as the time-critical aspect of the multi-agent system decision-making process. Twenty-four participants planned a multi-agent search and rescue mission, with the aid of an intelligent assistant. The results show that while the combined use of touch and speech was faster than speech alone, the single modality, touch, was still the most efficient. Future research should investigate additional input technologies.

Lisa R Le Vie↗

Exploring Multimodal Interactions in Human-Autonomy Teaming Using a Natural User Interface

The creation of a multimodal, natural user interface to facilitate multi-agent interaction is essential to establishing trust among human and machine teammates in multi-agent systems. Trust is being researched, along with trustworthiness, as a path to certification of autonomous systems by the Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR) project at NASA. The Autonomous Mission Experimental Logistics Interactive Assistant (AMELIA) is a natural user interface that enables multimodal interaction and is designed for rapid mission planning. AMELIA is an intelligent system that considers the user’s preferred communication strategies, as well as the time-critical aspect of the multi-agent system decision-making process. Twenty-four participants planned a multi-agent search and rescue mission, with the aid of an intelligent assistant. The results show that while the combined use of touch and speech was faster than speech alone, the single modality, touch, was still the most efficient. Future research should investigate additional input technologies.

Lisa Renee Le Vie↗

The role of artificial intelligence and expert systems in increasing STS operations productivity

Artificial Intelligence (AI) is discussed. A number of the computer technologies pioneered in the AI world can make significant contributions to increasing STS operations productivity. Application of expert systems, natural language, speech recognition, and other key technologies can reduce manpower while raising productivity. Many aspects of STS support lend themselves to this type of automation. The artificial intelligence section of the mission planning and analysis division has developed a number of functioning prototype systems which demonstrate the potential gains of applying AI technology.

Culbert, C.↗

i-SAIRAS '90; Proceedings of the International Symposium on Artificial Intelligence, Robotics and Automation in Space, Kobe, Japan, Nov. 18-20, 1990

The present conference on artificial intelligence (AI), robotics, and automation in space encompasses robot systems, lunar and planetary robots, advanced processing, expert systems, knowledge bases, issues of operation and management, manipulator control, and on-orbit service. Specific issues addressed include fundamental research in AI at NASA, the FTS dexterous telerobot, a target-capture experiment by a free-flying robot, the NASA Planetary Rover Program, the Katydid system for compiling KEE applications to Ada, and speech recognition for robots. Also addressed are a knowledge base for real-time diagnosis, a pilot-in-the-loop simulation of an orbital docking maneuver, intelligent perturbation algorithms for space scheduling optimization, a fuzzy control method for a space manipulator system, hyperredundant manipulator applications, robotic servicing of EOS instruments, and a summary of astronaut inputs on automation and robotics for the Space Station Freedom.

Source record↗

The adult literacy evaluator: An intelligent computer-aided training system for diagnosing adult illiterates

An important part of NASA's mission involves the secondary application of its technologies in the public and private sectors. One current application being developed is The Adult Literacy Evaluator, a simulation-based diagnostic tool designed to assess the operant literacy abilities of adults having difficulties in learning to read and write. Using ICAT system technology in addition to speech recognition, closed-captioned television (CCTV), live video and other state-of-the art graphics and storage capabilities, this project attempts to overcome the negative effects of adult literacy assessment by allowing the client to interact with an intelligent computer system which simulates real-life literacy activities and materials and which measures literacy performance in the actual context of its use. The specific objectives of the project are as follows: (1) To develop a simulation-based diagnostic tool to assess adults' prior knowledge about reading and writing processes in actual contexts of application; (2) to provide a profile of readers' strengths and weaknesses; and (3) to suggest instructional strategies and materials which can be used as a beginning point for remediation. In the first and developmental phase of the project, descriptions of literacy events and environments are being written and functional literacy documents analyzed for their components. Examples of literacy events and situations being considered included interactions with environmental print (e.g., billboards, street signs, commercial marquees, storefront logos, etc.), functional literacy materials (e.g., newspapers, magazines, telephone books, bills, receipts, etc.) and employment related communication (i.e., job descriptions, application forms, technical manuals, memorandums, newsletters, etc.). Each of these situations and materials is being analyzed for its literacy requirements in terms of written display (i.e., knowledge of printed forms and conventions), meaning demands (i.e., comprehension and word knowledge) and social situation. From these descriptions, scripts are being generated which define the interaction between the student, an on-screen guide and the simulated literacy environment. The proposed outcome of the Evaluator is a diagnostic profile which will present broad classifications of literacy behaviors across the major areas of metacognitive abilities, word recognition, vocabulary knowledge, comprehension and writing. From these classifications, suggestions for materials and strategies for instruction with which to begin corrective action will be made. The focus of the Literacy Evaluator will be essentially to provide an expert diagnosis and an interpretation of that assessment which then can be used by a human tutor to further design and individualize a remedial program as needed through the use of an authoring system.

Yaden, David B., Jr.↗

Intelligent Response and Interaction System (IRIS) - FY21 Closeout Report

In the second year, the IRIS team developed the components necessary for successful offline deployment of the IRIS services. This includes custom automated speech recognition training on NASA audio data, and in-house development and integration of online and offline conversational services. Lastly, the team worked on integrating the IRIS technology with stakeholders and projects that have a strong need for voice interaction.

Aly Shehata↗