Search NASA⌕ Search

SEARCH · Search NASA

Results for “Speech Intelligibility”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Sinusoidal transform coding

It has been shown that an analysis/synthesis system based on a sinusoidal representation of speech leads to synthetic speech that is essentially perceptually indistinguishable from the original. Strategies for coding the amplitudes, frequencies and phases of the sine waves have been developed that have led to a multirate coder operating at rates from 2400 to 9600 bps. The encoded speech is highly intelligible at all rates with a uniformly improving quality as the data rate is increased. A real-time fixed-point implementation has been developed using two ADSP2100 DSP chips. The methods used for coding and quantizing the sine-wave parameters for operation at the various frame rates are described.

Mcaulay, Robert J.↗

A comparison of two neural network schemes for navigation

Neural networks have been applied to tasks in several areas of artificial intelligence, including vision, speech, and language. Relatively little work has been done in the area of problem solving. Two approaches to path-finding are presented, both using neural network techniques. Both techniques require a training period. Training under the back propagation (BPL) method was accomplished by presenting representations of (current position, goal position) pairs as input and appropriate actions as output. The Hebbian/interactive activation (HIA) method uses the Hebbian rule to associate points that are nearby. A path to a goal is found by activating a representation of the goal in the network and processing until the current position is activated above some threshold level. BPL, using back-propagation learning, failed to learn, except in a very trivial fashion, that is equivalent to table lookup techniques. HIA, performed much better, and required storage of fewer weights. In drawing a comparison, it is important to note that back propagation techniques depend critically upon the forms of representation used, and can be sensitive to parameters in the simulations; hence the BPL technique may yet yield strong results.

Munro, Paul W.↗

A comparison of two neural network schemes for navigation

Neural networks have been applied to tasks in several areas of artificial intelligence, including vision, speech, and language. Relatively little work has been done in the area of problem solving. Two approaches to path-finding are presented, both using neural network techniques. Both techniques require a training period. Training under the back propagation (BPL) method was accomplished by presenting representations of current position, goal position pairs as input and appropriate actions as output. The Hebbian/interactive activation (HIA) method uses the Hebbian rule to associate points that are nearby. A path to a goal is found by activating a representation of the goal in the network and processing until the current position is activated above some threshold level. BPL, using back-propagation learning, failed to learn, except in a very trivial fashion, that is equivalent to table lookup techniques. HIA, performed much better, and required storage of fewer weights. In drawing a comparison, it is important to note that back propagation techniques depend critically upon the forms of representation used, and can be sensitive to parameters in the simulations; hence the BPL technique may yet yield strong results.

Munro, Paul↗

Considerations for Implementing Voice-Controlled Spacecraft Systems Through a Human-Centered Design Approach

As computational power and speech recognition algorithms improve, the consumer market will see better-performing speech recognition applications. The cell phone and Internet-related service industry have further enhanced speech recognition applications using artificial intelligence and statistical data-mining techniques. These improvements to speech recognition technology (SRT) may one day help astronauts on future deep space human missions that require control of complex spacecraft systems or spacesuit applications by voice. Though SRT and more advanced speech recognition techniques show promise, use of this technology for a space application such as vehicle/habitat/spacesuit requires careful considerations. There are still recognition challenges to overcome such as background noise, human speech variability, and task loading. However, implemented correctly, a voice-controlled spacecraft system (VCSS) can provide a useful, natural and efficient form of human-machine communications during complex tasks such as when the hands and eyes are busy or as an aid for vehicle situational awareness inquiries or collaborative human-robot tasks. This document provides considerations and guidance for the use of SRT in VCSS applications for space missions, specifically in command-and-control (C2) applications where the commanding is user-initiated. First, current SRT limitations as known at the time of this report are given. Then, highlights of SRT used in the space program provide the reader with a history of some of the human spaceflight applications and research. Next, an overview of the speech production process and the intrinsic variations of speech are provided. Finally, general guidance and considerations are given for development of a VCSS using a human-centered design approach for space applications that includes vocabulary selection and performance testing, as well as VCSS considerations for C2 dialogue management design, feedback, error handling, and evaluation/usability testing. Available from

Salazar, George A.↗

Development Considerations for Implementing a Voice-Controlled Spacecraft System

As computational power and speech recognition algorithms improve, the consumer market will see better-performing speech recognition applications. The cell phone and Internet-related service industry have further enhanced speech recognition applications using artificial intelligence and statistical data-mining techniques. These improvements to speech recognition technology (SRT) may one day help astronauts on future deep space human missions that require control of complex spacecraft systems or spacesuit applications by voice. Though SRT and more advanced speech recognition techniques show promise, use of this technology for a space application such as vehicle/habitat/spacesuit requires careful considerations. This paper provides considerations and guidance for the use of SRT in voice-controlled spacecraft systems (VCSS) applications for space missions, specifically in command-and-control (C2) applications where the commanding is user-initiated. First, current SRT limitations as known at the time of this report are given. Then, highlights of SRT used in the space program provide the reader with a history of some of the human spaceflight applications and research. Next, an overview of the speech production process and the intrinsic variations of speech are provided. Finally, general guidance and considerations are given for the development of a VCSS using a human-centered design approach for space applications that includes vocabulary selection and performance testing, as well as VCSS considerations for C2 dialogue management design, feedback, error handling, and evaluation/usability testing.

Salazar, George↗

Asynchronous sampling of speech with some vocoder experimental results

The method of asynchronously sampling speech is based upon the derivatives of the acoustical speech signal. The following results are apparent from experiments to date: (1) It is possible to represent speech by a string of pulses of uniform amplitude, where the only information contained in the string is the spacing of the pulses in time; (2) the string of pulses may be produced in a simple analog manner; (3) the first derivative of the original speech waveform is the most important for the encoding process; (4) the resulting pulse train can be utilized to control an acoustical signal production system to regenerate the intelligence of the original speech.

Babcock, M. L.↗

Universal Fourier Attack for Time Series

A wide variety of adversarial attacks have been proposed and explored using image and audio data. These attacks are notoriously easy to generate digitally when the attacker can directly manipulate the input to a model, but are much more difficult to implement in the real world. In this paper we present a universal, time invariant attack for general time series data such that the attack has a frequency spectrum primarily composed of the frequencies present in the original data. The universality of the attack makes it fast and easy to implement as no computation is required to add it to an input, while time invariance is useful for real world deployment. Additionally, the frequency constraint ensures the attack can withstand filtering defenses. We demonstrate the effectiveness of the attack on two different classification tasks through both digital and real world experiments, and show that the attack is robust against common transform-and-compare defense pipelines.

97 MATHEMATICS AND COMPUTING↗

Implementing Artificial Intelligence Behaviors in a Virtual World

In this paper, we will present a look at the current state of the art in human-computer interface technologies, including intelligent interactive agents, natural speech interaction and gestural based interfaces. We describe our use of these technologies to implement a cost effective, immersive experience on a public region in Second Life. We provision our Artificial Agents as a German Shepherd Dog avatar with an external rules engine controlling the behavior and movement. To interact with the avatar, we implemented a natural language and gesture system allowing the human avatars to use speech and physical gestures rather than interacting via a keyboard and mouse. The result is a system that allows multiple humans to interact naturally with AI avatars by playing games such as fetch with a flying disk and even practicing obedience exercises using voice and gesture, a natural seeming day in the park.

Krisler, Brian↗

The Next Wave: Humans, Computers, and Redefining Reality

The Augmented/Virtual Reality (AVR) Lab at KSC is dedicated to " exploration into the growing computer fields of Extended Reality and the Natural User Interface (it is) a proving ground for new technologies that can be integrated into future NASA projects and programs." The topics of Human Computer Interface, Human Computer Interaction, Augmented Reality, Virtual Reality, and Mixed Reality are defined; examples of work being done in these fields in the AVR Lab are given. Current new and future work in Computer Vision, Speech Recognition, and Artificial Intelligence are also outlined.

Extended Reality↗

xEMU Suit Integrated Audio Communications System: Ambient and EVA Pressure Testing System Performance

Testing across several airlock and EVA thermal and pressure scenarios has demonstrated that the Integrated Audio System of NASA’s Exploration Extravehicular Mobility Unit (xEMU) spacesuits transmits and receives intelligible audio communications without the use of a commcap or similar worn device. The xEMU audio system consists of internal loudspeakers and digital microphones (Integrated Communications System –ICS) combined with an adaptive Acoustic Echo Canceller (AEC), outbound voice operated transmission (VOX), and automatic gain control (AGC). Transducers are mounted in an “exploded commcap” configuration with helmet-attached speakers near the ears and three microphones positioned at the collar. The AGC removes inbound audio signals (e.g.,suit, Mission Control, Lander, C&W tones) from the outbound comms stream, reducing echo and feedback (squeal) in low-noise suit environments. The reduction of worn communication equipment increases crewmember comfort, range of movement, and situational awareness. However, test results also highlight the need for proper fan, duct, pump, and gas flow integration with suit acoustics and audio. Ductwork may serve as waveguides for various component and structure-borne noise. Sharply angled ducts can generate turbulent-flow noise. Gas flow from inlets above the crewmember’s head can generate noise when cascading over the faceplate and collar (or commcap) microphones. Sufficient acoustic noise levels (1) require increased gain to boost inbound audio, and (2) may distort signals resulting in AEC disruption or artifacts. Suit-noise levels decline with reduced pressure (density), but then elevated speech and audio effort/power become necessary. Whether the Integrated Audio System, commcap, or other device is used, suit acoustic noise can mask speech in outbound comms. This reduces intelligibility and requires other AECs/devices to suppress comms noise. Yet, adjusting a few components may yield significant improvement. We discuss xEMU audio functionality, demonstrate how acoustical treatment combined with inbound signal conditioning improved clarity during tests, and discuss future modifications.

xEMU↗

xEMU Suit Integrated Audio Communications System: Ambient and EVA Pressure Testing System Performance

Testing across several airlock and EVA thermal and pressure scenarios has demonstrated that the Integrated Audio System of NASA’s Exploration Extravehicular Mobility Unit (xEMU) spacesuits transmits and receives intelligible audio communications without the use of a commcap or similar worn device. The xEMU audio system consists of internal loudspeakers and digital microphones (Integrated Communications System –ICS) combined with an adaptive Acoustic Echo Canceller (AEC), outbound voice operated transmission (VOX), and automatic gain control (AGC). Transducers are mounted in an “exploded commcap” configuration with helmet-attached speakers near the ears and three microphones positioned at the collar. The AGC removes inbound audio signals (e.g.,suit, Mission Control, Lander, C&W tones) from the outbound comms stream, reducing echo and feedback (squeal) in low-noise suit environments. The reduction of worn communication equipment increases crewmember comfort, range of movement, and situational awareness. However, test results also highlight the need for proper fan, duct, pump, and gas flow integration with suit acoustics and audio. Ductwork may serve as waveguides for various component and structure-borne noise. Sharply angled ducts can generate turbulent-flow noise. Gas flow from inlets above the crewmember’s head can generate noise when cascading over the faceplate and collar (or commcap) microphones. Sufficient acoustic noise levels (1) require increased gain to boost inbound audio, and (2) may distort signals resulting in AEC disruption or artifacts. Suit-noise levels decline with reduced pressure (density), but then elevated speech and audio effort/power become necessary. Whether the Integrated Audio System, commcap, or other device is used, suit acoustic noise can mask speech in outbound comms. This reduces intelligibility and requires other AECs/devices to suppress comms noise. Yet, adjusting a few components may yield significant improvement. We discuss xEMU audio functionality, demonstrate how acoustical treatment combined with inbound signal conditioning improved clarity during tests, and discuss future modifications.

xEMU↗

Multipath/RFI/modulation study for DRSS-RFI problem: Voice coding and intelligibility testing for a satellite-based air traffic control system

Analog and digital voice coding techniques for application to an L-band satellite-basedair traffic control (ATC) system for over ocean deployment are examined. In addition to performance, the techniques are compared on the basis of cost, size, weight, power consumption, availability, reliability, and multiplexing features. Candidate systems are chosen on the bases of minimum required RF bandwidth and received carrier-to-noise density ratios. A detailed survey of automated and nonautomated intelligibility testing methods and devices is presented and comparisons given. Subjective evaluation of speech system by preference tests is considered. Conclusion and recommendations are developed regarding the selection of the voice system. Likewise, conclusions and recommendations are developed for the appropriate use of intelligibility tests, speech quality measurements, and preference tests with the framework of the proposed ATC system.

Birch, J. N.↗

Comparison of voice types for helicopter voice warning systems

Three related studies were conducted to compare different types of human voice warnings. In the first study, a comparison of three LPC-encoded voices, human female, human male, and phoneme-synthesized, by the criteria of pilot flight task performance showed no differences due to the voice type. In the second study, pilots' preferences were investigated, by comparing preference for direct synthesized speech to the LPC-encoded human female speech and to LPC-encoded synthesized speech. Most pilots were found to prefer direct synthesized speech over both LPC-encoded human female speech and the LPC-encoded synthesized speech. In the third study, phonetically balanced (PB) words heard in simulated helicopter noise were used to compare the intelligibility of direct synthesized and LPC-encoded phoneme-synthesized speech types. PB word intelligibility was found to be better for direct synthesized speech than for the LPC-encodes synthesized speech.

Simpson, C. A.↗

Quantization noise in digital speech

The amount of quantization noise generated in a digital-to-analog converter is dependent on the number of bits or quantization levels used to digitize the analog signal in the analog-to-digital converter. The minimum number of quantization levels and the minimum sample rate were derived for a digital voice channel. A sample rate of 6000 samples per second and lowpass filters with a 3 db cutoff of 2400 Hz are required for 100 percent sentence intelligibility. Consonant sounds are the first speech components to be degraded by quantization noise. A compression amplifier can be used to increase the weighting of the consonant sound amplitudes in the analog-to-digital converter. An expansion network must be installed at the output of the digital-to-analog converter to restore the original weighting of the consonant sounds. This technique results in 100 percent sentence intelligibility for a sample rate of 5000 samples per second, eight quantization levels, and lowpass filters with a 3 db cutoff of 2000 Hz.

Schmidt, O. L.↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Measurement of speech levels in the presence of time varying background noise

Short-term speech level measurements which could be used to note changes in vocal effort in a time varying noise environment were studied. Knowing the changes in speech level would in turn allow prediction of intelligibility in the presence of aircraft flyover noise. Tests indicated that it is possible to use two second samples of speech to estimate long term root mean square speech levels. Other tests were also performed in which people read out loud during aircraft flyover noise. Results of these tests indicate that people do indeed raise their voice during flyovers at a rate of about 3-1/2 dB for each 10 dB increase in background level. This finding is in agreement with other tests of speech levels in the presence of steady state background noise.

Pearsons, K. S.↗

Deep Generative Models in Energy System Applications: Review, Challenges, and Future Directions

In recent years, with the advent of mature machine learning products like ChatGPT, Stable Diffusion, and Sora, the world has witnessed tremendous changes driven by the rapid development of generative artificial intelligence (GAI). Beyond applications in text, speech, image, and video creation, deep generative models (DGMs) underpinning these cutting-edge technologies have also been employed by domain researchers to address scientific and engineering challenges. This paper aims to fill a gap in the research community by providing a systematic review of how DGMs have been utilized in energy system applications. After introducing four most popular DGMs, we review and categorize 196 research articles into five focus areas: data generation, forecasting, situational awareness, modeling, and optimal decision-making. Through this classification, we uncover trends in how DGMs are employed for each type of problem, highlighting GAI techniques that contribute to breakthroughs over traditional methods. We discuss limitations in existing literature, engineering challenges, and propose future directions, all tailored to the unique nature of problems in energy system engineering. Our goal is to offer insights for energy system domain researchers, providing a comprehensive view of existing studies and potential future opportunities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Call sign intelligibility improvement using a spatial auditory display

A spatial auditory display was used to convolve speech stimuli, consisting of 130 different call signs used in the communications protocol of NASA's John F. Kennedy Space Center, to different virtual auditory positions. An adaptive staircase method was used to determine intelligibility levels of the signal against diotic speech babble, with spatial positions at 30 deg azimuth increments. Non-individualized, minimum-phase approximations of head-related transfer functions were used. The results showed a maximal intelligibility improvement of about 6 dB when the signal was spatialized to 60 deg or 90 deg azimuth positions.

Begault, Durand R.↗